i don't like to bring my professional life onto tumblr but i am an interdisciplinary computer scientist that has worked within these institutions and has worked extensively with models like this firsthand - my greatest concern throughout all of this is implementations just like this.
this is something thats been brushed off (of course) by my field and i've butted heads with other researchers in my field about the push to integrate these models with no regard to how unequitable they are. before the stanford debacle was made public, i had already taken great issue with the NER (Named Entity Recognition) models that stanford had been pushing, whether it be internal implements or publications released by the university. i have pushed these models to their limits and have found hundreds of times over that they will never, ever, ever be equitable in any sense of the word. i have developed bias audit suites to stress test these models to ensure that they are compliant with the ethical frameworks that we, as researchers, strive to model.
here's a bit of what i mean. the following screenshot is coming from a manuscript of mine that has yet to be published:
the model that i put forth above met the pictured metrics after i had done the barest minimum for a named entity recognition model using an augmented dataset specifically constructed to increase the representation of names from African, Afro-Caribbean, Black American, East African, and West African naming traditions.
the augmented dataset was constructed through a combination of census mining, public domain biographical text, and synthetic augmentation using attested name lists drawn from demographic and sociological research on naming practices in Black communities in the united states, the caribbean, and the african continent. this is the barest minimum that we should undertake and it only scratches the surface of what is necessary to make our machine learning implementations less catastrophic, and still fails to meet expectations of what a truly equitable model should be. the inclusion of synthetic augmentation is acknowledged as a methodological compromise because, ideally, training data would consist of naturally occurring text in which names from all represented traditions appear as entities in realistic contexts.
i can say that i have sat in rooms where these models were trained for years and no entity, be it an academic entity or a corporate entity, has ever given it a fraction of this much thought. i have destroyed personal, consumer-grade hardware that i have paid for and continue to pay for out of my own pocket in order to continue these audits.
i have raised formal complaints and submitted requests for investigation with publishers and institutions about problematic pursuits that cannot pass bias audits and the peer reviewers of these publications seldom review NER models in this much detail. my Black colleagues have done the same and rarely, if ever, receive responses. i have seen a Black colleague defer their complaints and words to a white scientist, as the exact same words have only been taken seriously when the messenger demographic is the same as the offender. i have seen these labs receive funding, stating specific west african benefit in their grant requests, only for no occupant of the lab to know what the words 'Yoruba, Igbo, Akan, and Wolof' even are when prompted. the funding is never revoked and no one is ever disciplined save for the whistleblowers. NER training corpora excludes all of these groups knowing that there will never be consequences for doing so and that the systemic racism perpetuated will be rewarded every single time by bolstering the careers of the researchers that enable it.
i have spent hundreds if not thousands of hours developing equitable and ethical machine learning implementations for research, entirely unfunded, as you can guess where the grant money is deferred to. i have lost significant career advancement opportunities because, in order to hold these positions, you must relinquish your willingness to be an ally of any person of color. i have spent years watching interns who i believed to have great moral character become perpetrators of racial violence through greed and selfishness. now imagine what happens to the researcher that isn't a white man like myself - for many, there is little to no chance to seize the privilege of being able to develop these implementations. there is no chance to enter these rooms, enter these discussions, reap the benefits of the mobility that this specialized education gives to you, or have any say on what happens to your demographic in the model architecture. those who make it will be pressured into submission and those who fight against it will be suppressed somehow.
it is more important than ever for Black computer scientists to be uplifted and for their research to be funded and made visible - it is these same scientists being identified by these models as "out-of-vocabulary" or non-entities, while dozens of less qualified peers are identified correctly.
my worst fear is NER automation or screening making it to the grant distribution process: it will be catastrophic for research as the distribution of grants is once again tailored to the white ideal and sole beneficiary. this technical failure is not neutral and is unforgivable in the "logical" field of computer science.