CMMLU: Chinese Massive Multitask Language Understanding — A Native Chinese LLM Benchmark Built for Domestic Cultural & Linguistic Scenarios
With the rapid advancement of large language models, a critical question has come to the fore: what is the real capability of these supposedly omniscient AI systems within Chinese linguistic contexts? Especially when confronted with knowledge and expressions unique to Chinese culture, models trained predominantly on English corpora consistently display glaring weaknesses. To fill this evaluation gap, research teams from Shanghai Jiao Tong University, Tsinghua University and other institutions jointly developed CMMLU (Chinese Massive Multitask Language Understanding), a comprehensive benchmark tailor-made for Chinese large language models.
Standout Feature: Full Coverage & Deep Localization for Chinese Culture
Unlike many benchmark datasets simply translated from English originals, CMMLU test items are designed from the ground up to account for the unique logic and cultural nuances of the Chinese language. The benchmark spans 67 distinct subject categories, covering three tiers:
- Foundational disciplines: basic mathematics, physics, chemistry;
- Advanced professional fields: law, economics, clinical medicine;
- Culture-specific Chinese content: China’s traffic regulations, Chinese history, classical Chinese literature, and more.
Many test questions hold answers only valid within China’s cultural and institutional context, which do not apply to other regions or languages. To achieve high CMMLU scores, a model cannot merely rely on cross-lingual translation; it must genuinely grasp China’s unique knowledge system and native modes of reasoning.
Solid Dataset Scale & Standardized Four-Option Multiple-Choice Format
CMMLU contains 11,582 formal test questions plus 335 development set samples, all structured as four-option single-answer multiple-choice items. Example biology question: Two cell types from the same species each secrete a distinct protein. The two proteins share identical amino acid composition ratios but differ in amino acid sequence. What accounts for this difference? A. Different types of tRNA B. One codon encoding different amino acids C. Different base sequences of mRNA D. Differently structured ribosomes Correct answer: C
Uniform question formatting and unambiguous objective answers enable fully automated, standardized quantitative evaluation, with scores from different models directly comparable.
Official Leaderboard: Qwen2 72B Instruct Breaks the 90% Accuracy Threshold
On the latest CMMLU ranking, Alibaba’s Tongyi Qwen series delivers standout performance. The Qwen2 72B Instruct model takes first place with an accuracy of 90.1%, the first model ever to surpass the 90% milestone on this benchmark. It is followed by LongCat-Flash-Chat (84.3%) and LongCat-Flash-Lite (82.5%). These metrics demonstrate that state-of-the-art Chinese LLMs have reached a high standard of knowledge comprehension and logical reasoning, while providing clear quantifiable metrics to guide ongoing technical iteration.
Widely Recognized Academic Value & Open Non-Commercial License
The research paper for CMMLU thoroughly documents dataset construction workflows and standardized evaluation protocols. Released under the CC BY-NC-SA 4.0 open license, the dataset is freely downloadable by researchers worldwide for non-commercial academic work. It has become an indispensable authoritative reference for three core user groups: AI companies conducting internal model capability testing, academic institutions researching Chinese natural language processing, and developers comparing the native Chinese proficiency of competing large models.