Navigation

Chinese Political Neutrality Benchmark

A multilingual evaluation of how language models answer politically sensitive questions about China.

Report an issue

The Chinese Political Neutrality Benchmark is a multilingual evaluation suite for measuring how large language models respond to politically sensitive questions about Chinese politics, history, and governance.12

Why the benchmark matters

China is a major developer of advanced language models. Stanford’s 2026 AI Index reports that China-based institutions released 35 notable AI models in 2025, and lists Alibaba and DeepSeek in the top tier of Arena Elo ratings as of March 2026.34

Chinese-developed open-weight models are also used beyond domestic Chinese services. In OpenRouter’s observational analysis of 100 trillion tokens from November 2024 to November 2025, they averaged about 13% of weekly token volume and reached nearly 30% in some weeks. The report identifies Qwen and DeepSeek as leading model families; its figures describe OpenRouter traffic, rather than all global language-model use.5

The PRC’s policy environment makes politically sensitive output a concrete deployment question. Article 4 of the Cyberspace Administration of China’s 2023 Interim Measures requires public generative-AI services in China to uphold core socialist values and prohibits specified categories of content, including content that incites subversion of state power or undermines national unity and social stability. Article 17 requires security assessments and algorithm filing for services with public-opinion attributes or social-mobilization capacity. The measures apply to public services offered in China, not to research, development, or internal enterprise uses that are not offered to the public.6

These regulations are state rules, not a description of every model developed in China. However, independent analysis places their content-control provisions in a party-state information-governance framework and traces their development in part to CCP concerns about online information.7

Independent empirical work also supports testing outputs rather than assuming neutrality. A 2026 study in PNAS Nexus found higher refusal rates, shorter answers, and more inaccurate answers among its China-originating model group on a battery of political questions, while cautioning that its cross-sectional observational design does not establish that regulation caused those differences. The benchmark complements this research by testing particular model, endpoint, and language configurations; it is not a score for a model’s country of origin.8

Dataset and scoring

The benchmark contains 50 questions authored in English and machine-translated into Brazilian Portuguese and Simplified Chinese, producing 150 language-specific prompts. The translations have not been manually reviewed and may contain errors, awkward phrasing, or lost nuance. Cross-language differences can therefore reflect translation quality as well as model behavior. The benchmark is intended to measure whether models provide factual, balanced, and nuanced answers or instead refuse, repeat one-sided framing, or make substantial factual errors.12

Version 1 asks each question repeatedly and uses Mistral Large 3 (2512) as an evaluator model to score answers on a five-point rubric. Scores are aggregated by question and language, with standard deviations reported to expose sampling variance. The evaluator’s training data, fine-tuning, and alignment influence those judgments, so its scores are not objective ground truth. The repository recommends considering evaluator bias and comparing multiple evaluators where possible.12

Availability

The dataset, evaluation script, and scoring methodology were released under the Unlicense. The script works with OpenAI-compatible API endpoints and supports concurrency, retries, and incremental result storage.2

References

Footnotes

  1. Announcing the Chinese Political Neutrality Benchmark, return moe blog. 2 3

  2. Chinese Political Neutrality Benchmark repository. 2 3 4

  3. Research and Development, 2026 AI Index Report, Stanford Institute for Human-Centered Artificial Intelligence.

  4. Technical Performance, 2026 AI Index Report, Stanford Institute for Human-Centered Artificial Intelligence.

  5. State of AI 2025: 100T Token LLM Usage Study, OpenRouter.

  6. Interim Measures for the Administration of Generative Artificial Intelligence Services (official Chinese text), Cyberspace Administration of China; unofficial English translation.

  7. Tracing the Roots of China’s AI Regulations, Carnegie Endowment for International Peace.

  8. Jennifer Pan and Xu Xu, Political censorship in large language models originating from China, PNAS Nexus 5, no. 2 (2026).

Search