spot_imgspot_img

AI safety is designed in the West, and failing users everywhere

Last month, OpenAI became the first major artificial intelligence company to voluntarily pause training on a model because of safety concerns. The unprecedented move came weeks after its models broke free during a test and hacked other websites, sparking concern about losing control of AI systems, as Anthropic and Meta reported similar incidents.

“We care very deeply about AI safety,” OpenAI chief executive Sam Altman said on X. “Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment.”

AI researchers had been calling for a slower pace for development and more safety measures even before these incidents. Generative AI systems are solving complex mathematical problems and helping develop highly specialized drugs in the U.S. and other Western nations. But they are failing at basic tasks elsewhere because trust and safety teams are largely concentrated in Silicon Valley, and do not reflect the concerns of countries with different languages and cultural contexts.

The gaps have real consequences: Queries related to health are among the most common uses of AI chatbots worldwide, yet in many African and Asian nations, even multilingual AI tools make errors that can affect diagnoses and treatment decisions. More than two-thirds of chatbots do not adequately account for dialects or recognize urgency cues, according to a review in India.

At the heart of the issue “is the question of who gets to define what counts as a safety problem in the first place.”Elizabeth Orembo, a fellow at Research ICT Africa

The global majority “still remains at the margins of the larger AI safety discourse,” Urvashi Aneja, founder of research organization Digital Futures Lab, who is working on a report for the United Nations on AI safety in developing nations, told Rest of World. “The frameworks being built to evaluate AI systems, the standards that govern them, and the institutions that oversee them have largely been designed in, and for, a small set of high-income countries.”

Trust and safety is an umbrella term for the teams at tech companies whose job is to ensure that users are protected from harmful online experiences. There is no universal trust and safety standard; each company has its own framework — based on its values and principles, and enforcement methods.

Developing nations are adopting AI at a slower pace than wealthier nations, yet the risks fall “disproportionately” on them because of inadequate resources, limited domestic AI infrastructure, and dependence on foreign technologies, the U.N. said in a recent report.

At the heart of the issue “is the question of who gets to define what counts as a safety problem in the first place,” Elizabeth Orembo, a fellow at Research ICT Africa, a think tank, told Rest of World. Big tech firms tend to focus on model risks such as deception, autonomous behavior, cyber capabilities, and aiding bioweapons, she said. They do not pay much attention to deployment risks including discrimination, exclusion, surveillance, language failures, and the inability of affected communities to seek remediation.

Life-threatening consequences

Trust and safety teams run evaluations that assume reliable electricity and internet connectivity, functioning courts, robust data protection laws, formal labor markets, and a press and civil society that report failures. So “a model can pass every frontier safety evaluation and still produce unsafe outcomes when deployed,” Orembo said.

Natural language processing in healthcare in Africa showed cultural and linguistic bias, poor adaptation to medical contexts, and translation errors, researchers found. For example, in Tigrinya, spoken by about 9 million people in Eritrea and northern Ethiopia, machine translation rendered smallpox as syphilis, gonorrhea as diabetes and “you have been given intravenous antibiotics” as “you have been given intravenous insecticides.”

Such mistranslations “can be life-threatening,” Orembo said.

On a recent AI safety index from the Future of Life Institute that evaluated nine leading companies, Anthropic, OpenAI, and Meta had the highest scores on metrics such as risk assessment, current harms, existential safety, and governance and accountability. DeepSeek, xAI, and Mistral had the lowest scores.

But “even industry leaders … are retreating from prior commitments, despite calling publicly for a pause,” Future of Life Institute, a nonprofit that researches AI risks, noted. This has “undermined safety frameworks across the board.”

In low- and middle-income countries, the consequences are compounded, and users feel the effects “immediately” as they can affect access to wages and essential services, the U.N. Development Programme said in a recent report.

There is growing evidence of these failures, including AI-powered facial recognition and ID systems that lead to denial of wages, meals, or school attendance, and tools that misidentify crops or mistranslate local terms.

The guardrails “may work well in English, but fail or are easily circumvented in low-resource languages.”Dhanaraj Thakur, director of the Fair Technology Initiative at the George Washington University Law School

Language failures are inevitable, as data sets to train large language models are predominantly in English and other widely spoken Western languages, and LLMs display poor translation and higher hallucination rates in low-resource languages, Dhanaraj Thakur, director of the Fair Technology Initiative at the George Washington University Law School, told Rest of World.

The guardrails “may work well in English, but fail or are easily circumvented in low-resource languages,” he said. “The result is that users of these models that speak English or other high-resource languages end up being safer than those that speak low-resource languages, a new kind of AI divide.”

A critical moment

Governments are trying to address these concerns. At the inaugural AI summit in 2023 in the U.K., 28 countries signed the Bletchley Declaration, a voluntary commitment to identify and respond to existential risks tied to AI. The India AI summit earlier this year also addressed safety, while China — which is regulating AI more aggressively — in July proposed mechanisms to manage AI risks in developing nations.

For these countries, there is no time to lose, Aneja said.

“There is a lot of optimism about AI in these countries now, unlike the backlash that you see in the West,” she said. “But if governments don’t invest in the safety infrastructure now, public trust will erode, and the opportunity to leverage benefits from AI will go away, and then we’re looking at deepening inequality.”

Recently, more than 1,300 employees at AI companies including Anthropic, Meta AI, OpenAI, and Google DeepMind wrote an open letter saying the industry, the government, and society “may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight” of AI systems.

For Sumiya Khan, a 16-year-old in New Delhi, this is sorely needed. Experiencing constant fatigue and dizziness, she turned to ChatGPT for advice in Hindi, like she had done before. The chatbot chalked up her symptoms to stress and poor sleep. But when the symptoms persisted, Khan visited a doctor, who diagnosed iron-deficiency anemia, which is prevalent in low- and middle-income countries. Left untreated, it can lead to irregular heart rhythm and, in extreme cases, heart failure.

“We trusted the AI because it sounded so convincing,” Mehnaz Begum, her mother, told Rest of World. “That made us wait longer than we should have to see a doctor.”

Additional reporting by Sajid Raina in New Delhi.

0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments

Popular Articles

0
Would love your thoughts, please comment.x
()
x