
Regulators and researchers spent the past year warning that Chinese AI models like DeepSeek come with Beijing’s censorship baked in, but a growing body of research suggests American AI models may not be entirely immune from it.
A peer-reviewed study published in Nature found evidence that Chinese state-controlled media makes its way into AI training data and can influence how models answer questions about China. Meta’s independent Oversight Board found models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse requests to criticize governments in countries that restrict political speech than those in freer countries—raising questions about whether the rules of authoritarian information environments are migrating into AI products used around the world.
“This kind of influence is concerning because—like covert information operations—it severs information and opinion from their source, effectively laundering government-manipulated content into ostensibly objective text,” the Nature researchers wrote.
The evidence taken together suggests the effects of China’s tightly controlled information system may not stop at Chinese-built AI, though neither shows that Beijing deliberately manipulated OpenAI, Anthropic, Google or Meta.
It does point to a deeper problem for an industry that markets its models as politically neutral. Anthropic, for example, has touted efforts to make Claude politically “even-handed,” reporting a 94% score on its own evaluation.
How Chinese propaganda seeps into training data
The Nature paper researchers identified more than three million Chinese-language documents in the open-source training dataset CulturaX–used to train and improve LLMs— and built a “multi-part case study on China’s media” that particularly focused on political subjects. They also found signs that the models had encountered the material during training: Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4 and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%.
The researchers then tested what that kind of material could do to a model. After further training Meta’s open-weight Llama 2 13B—picked because it had very little to zero Chinese state media in its training data—on just 6,400 Chinese state-scripted news examples, the model produced a more Beijing-friendly answer than the baseline model nearly 80% of the time.
At higher levels of additional training, the difference became stark. After 64,000 state-scripted examples, the retrained model was asked whether China is an autocracy. The baseline model said it was. The state-scripted version instead described China as democratic and invoked the Chinese Communist Party’s concept of “people’s democracy.”
The researchers could not run the same training-data experiment on proprietary systems such as OpenAI’s and Anthropic’s, whose training processes are largely opaque. Instead, they asked the same political questions in Chinese and English to the models and compared the answers.
The gap was substantial. The Chinese-language response was rated as more favorable to Chinese leaders and institutions 68.8% of the time for Claude Sonnet, 88.2% for Claude Opus, 72.6% for GPT-3.5 and 84% for GPT-4o.
The effect was not confined to China. In another audit involving 6,051 prompts across 37 countries, the researchers found that countries with lower levels of press freedom tended to receive more favorable descriptions from GPT-3.5, GPT-4o, Claude Opus and Claude Sonnet when the models were queried in the country’s dominant language rather than in English.
“By disguising the source of the influence and incentives of the state, we fear that LLMs may have the potential to further increase the subtlety and persuasive power of state media control,” the researchers wrote.
The censorship problem goes beyond training data
The Oversight Board report found a different manifestation of the broader problem: American AI models sometimes behaved as though political restrictions from authoritarian countries applied even to users outside those countries, an issue the report calls “censorship-by-proxy.”
Researchers tested 10 commercial models from Anthropic, DeepSeek, Google, Meta, OpenAI and xAI. They used identical political prompts involving five countries with restrictive speech laws—China, Saudi Arabia, Thailand, Turkey and Cambodia—and five relatively permissive countries, including the U.S., U.K., Japan, Taiwan and Chile. The tests were conducted from Australia.
Across the 10 models, the average refusal rate for requests to produce political criticism was 34% in restrictive countries, compared with 14% in freer ones. But there was wide variation by model: Gemini 3 Flash and Grok 4 Fast, for example, did not refuse any of the political-material requests regardless of jurisdiction.
Some of the clearest disparities appeared in Anthropic’s Claude Sonnet 4. It refused all five protest flyer requests involving Xi Jinping, Saudi Crown Prince Mohammed bin Salman and Thailand’s King Vajiralongkorn, while producing all five requested flyers criticizing President Donald Trump and King Charles III. It also complied four out of five times for Chile’s then-president and three out of five times for Japan’s then-prime minister.
Google’s Gemini 3 Pro showed a similar pattern, but it refused three of five involving Xi, four of five involving Mohammed bin Salman and Thailand’s king, and all five involving Cambodia’s king. In nearly all of the Thailand and Cambodia refusals, the model’s reasoning invoked criminal restrictions or lèse-majesté laws.
Meta’s open-weight model Llama 4 Maverick (different from its new proprietary system Muse Spark) also complied with every flyer request involving Trump, King Charles, Japan, Chile, Taiwan and Turkey’s Recep Tayyip Erdogan, but refused all five involving Xi, Thailand’s king and Cambodia’s king. In one Xi refusal, it said criticism of government leaders could be “sensitive or illegal” in China.
“These results show that there is a real and concerning risk that foundation models could be reflecting and further entrenching the restrictive speech norms of repressive regimes,” the researchers wrote.
Meta declined to comment. Anthropic, OpenAI and Google did not respond to Fortune’s requests for comment.











