A benchmark for one specific failure

Aleph Alpha, the German AI company, has published a detection benchmark aimed at a narrow question: how often does an open-weight model answer a politically sensitive question about China in a way that reproduces the Chinese government’s framing? The work, written up by Bastian Boll and published on 28 September, uses 967 prompts across 56 manually selected topics.

Responses were classified by an LLM judge — GPT-OSS 120B — into three buckets: answers echoing Communist Party framing, balanced answers, and refusals. A second set of 240 prompts that never mention China was used to test whether the behaviour spills into unrelated questions.

The topics cluster into three themes: state security and leadership, including references to Xi Jinping; territorial questions covering Taiwan, Hong Kong and Xinjiang; and ideological concepts such as press freedom, human rights and democratic elections.

Assorted white ceramic pieces on a studio shelf
The benchmark uses 967 prompts across 56 manually selected topics. Illustrative photograph. Anastasia Shuraeva · pexels · Pexels License

The results

Six Chinese open-weight models were tested: Qwen 3.6 35B-A3B and Qwen 3.8 2.4T-A95B, DeepSeek R1 0528 and V4 Pro 0813, and Kimi K2.5 and K3. Non-Chinese baselines were Claude Sonnet 5, Mistral Small 2603 and Nvidia’s Nemotron Cascade 2.

Qwen 3.6 answered 17% of prompts in a balanced way and echoed Communist Party framing on 80% of them. Claude Sonnet 5 answered 70% in a balanced way, with 2.8% showing alignment. Across the Chinese models the balanced share ran between 17% and 41%, according to The Decoder’s summary of the work; the remainder repeated state framing, deflected or declined to answer.

None of this is surprising on its own. Chinese generative AI services operate under rules requiring content to uphold core socialist values, and the models are shaped accordingly. The benchmark’s contribution is a number rather than an impression.

Rows of illuminated server status lights in a rack
Aleph Alpha traced one Western model's results to its fine-tuning data. Illustrative photograph. panumas nikhomkhai · pexels · Pexels License

The result that matters more

Nvidia’s Nemotron Cascade 2 was flagged on 17% of prompts — a Western model, built by an American company, producing the same framing at a rate comparable to the lowest-scoring Chinese one.

Aleph Alpha’s explanation is contamination rather than intent. It reports finding roughly 3,500 rows out of 9.3 million in the model’s chat fine-tuning data that carried Communist Party talking points, likely because that supervised fine-tuning data was itself generated using Chinese models.

That is the finding with consequences beyond this benchmark. Synthetic fine-tuning data generated by whatever model is cheapest and most capable is now standard practice across the industry, and it carries whatever the generating model was shaped to believe. A fraction of a percent of a training set was enough to move a measurable share of answers.

What to watch is whether anyone starts auditing synthetic training data for inherited positions as a matter of routine. The tooling to do it clearly exists, because Aleph Alpha just did it to someone else’s model.