Archived
[...]
On a benchmark of 967 politically sensitive prompts, six Chinese models answered only 17 to 41 percent of prompts in a balanced way, a new study finds.
More surprisingly, NVIDIA’s Nemotron Cascade 2 was flagged on 17 percent of prompts. Its published SFT data was generated largely with DeepSeek and Qwen, and here we found the likely reason for the model’s observed behavior: roughly 3.5k of its 9.3M chat rows carrying Chinese Communist Party (CCP) talking points.
From this data, we conclude that sovereign models need three things: screening of training data for Chinese political content, targeted alignment data that sets the intended behavior, and evaluation against benchmarks like the one presented here. We have adopted these measures at Aleph Alpha.
[...]
A recent report by LatticeFlow AI found that, as Chinese models grow larger, their political leaning grows stronger, creating a wider gap to Western norms.
[...]
Although Chinese state influence has been reported, users of Chinese models may not notice it day-to-day. By probing these models on a large variety of topics, we have found that state-directed alignment is centered around the following themes.
[...]
Chinese models will express their political alignment in a number of ways, including by making doctrinal assertions in the assistant’s voice, referring to Chinese law as a reason to refuse, denying documented events, answering evasively or steering the user toward Chinese state media. When prompted about (lawful) political advocacy around these topics, we have also seen Chinese models propose courses of action that soften the user’s requests when they conflict with Chinese state positions.
[...]
The report evaluated Chinese LLM's responses to three topics:
- Direct mentions of top political leaders, the Chinese Communist Party’s legitimacy, and historical taboos like the Tiananmen Square protests
- Taiwan, Hong Kong, the South China Sea, and regions like Xinjiang and Tibet.
- Topics like press freedom, freedom of assembly, human rights, democratic elections, and minority rights.
When prompted about these topics, Chinese models will express their political alignment in a number of ways, including by making doctrinal assertions in the assistant’s voice, referring to Chinese law as a reason to refuse, denying documented events, answering evasively or steering the user toward Chinese state media. When prompted about (lawful) political advocacy around these topics, we have also seen Chinese models propose courses of action that soften the user’s requests when they conflict with Chinese state positions.
[...]
Overall, all evaluated Chinese models show strong alignment with CCP positions compared to the baselines trained outside of China, either by asserting doctrine or by refusing to answer.
[...]
Chinese state perspectives in non-Chinese LLMs
It is perhaps unsurprising that Chinese state narratives are produced by Chinese LLMs, since the regulation mentioned above explicitly demands political alignment. More surprising is the fact that, using the detection benchmark developed above, we found Chinese doctrinal assertions in responses from NVIDIA’s Nemotron Cascade 2.
[...]
NVIDIA publishes the full supervised finetuning (SFT) data that was used to train Cascade 2, and we conclude this to be the likely source: chat data responses were generated by Chinese models, including DeepSeek and Qwen. Since these datasets purposefully contain a wide variety of prompts, some fraction of them is also politically sensitive, triggering Chinese models to give the party line.
[...]
Still, CCP talking points leaking into Nemotron Cascade 2 makes clear that active care is required to safeguard model sovereignty from Chinese influence [...] The trend is clear: a significant fraction of the future internet will be generated by Chinese LLMs. Therein lies a risk of the continued world-wide adoption of Chinese LLMs: the coverage of sensitive topics on the internet will shift over time and with it the distribution of pre-training data.
[...]