As Large Model Technology Booms, AGI-Eval Delivers Authoritative, Human-Centric AI Benchmarking Co-Launched by Top Chinese Universities & Research Institutes
Hundreds of AI large models have emerged globally amid rapid technical advancement, yet researchers and developers face a persistent core challenge: how to fairly rank model performance and quantify the gap between AI and human cognitive capabilities. To resolve this pain point, top-tier universities including Shanghai Jiao Tong University, Tongji University, East China Normal University, alongside DataWhale and other leading institutions, jointly built AGI-Eval — an LLM evaluation community guided by the mission: Powered by benchmarking, make AI a better partner for humanity.
Unique Positioning: Benchmarking Built on Standard Human Cognitive Examinations
Unlike conventional evaluation systems that only measure isolated technical metrics, AGI-Eval assesses general artificial intelligence via high-stakes official human admissions tests, qualification exams and advanced competitive contests tailored for human test-takers. Its test bank covers classic selective assessments such as the LSAT law school entrance exam, China’s National College Entrance Examination (Gaokao), U.S. SAT, mathematics Olympiads and bar qualification exams.
This paradigm directly measures how well AI aligns with human decision-making and cognitive logic, accurately reflecting real-world practicality and reliability. To deliver comprehensive cross-cultural assessment, AGI-Eval integrates bilingual Chinese-English tasks, capturing authentic model performance across multilingual and cross-cultural scenarios.
Independent Third-Party Rigorous Benchmark Reports Released Throughout 2025
As an impartial third-party evaluation platform, AGI-Eval published dozens of thorough, objective, in-depth benchmark reports in 2025. Its curated annual report series systematically tests mainstream models including GPT-4o, DeepSeek, Qwen3, Claude, StepAI and Doubao.
- Specialized text-to-image benchmark for GPT-4o: Scoring across dimensions of text-image consistency, visual quality, commonsense reasoning and structured generation confirmed GPT-4o took the overall top spot with a composite score of 4.41, far ahead of second-place Dreamina 2.1 at 4.01. It delivered standout performance on structured tasks such as character rendering and chart drafting.
- Double-blind real-time voice interaction cross-test covering 8 leading products: Based on 1,624 authentic voice dialogue samples and assessments from 480 human testers, domestic models StepAI (0.64) and Doubao (0.63) outperformed GPT-4o (0.60) in overall conversational fluency.
Industry-Leading Innovative Modular Open-Source Evaluation Framework (Released November 2025)
AGI-Eval open-sourced its internal benchmarking framework with a core design philosophy: Evaluation is not a fixed pipeline, but a pluggable, extensible system. Built on a plug-in architecture, it supports local single-machine debugging, multi-process parallel execution and flexible concurrency adjustment based on hardware resources. Every link from data preprocessing to metric calculation can be packaged as independent plug-ins for free combination and expansion without modifying the core main framework. Native built-in web reporting tools support metric statistics, horizontal model comparison and error sample review. To guarantee fully reproducible evaluation results, the team fine-tuned a dedicated scoring judge model AGI-Eval-OA-Judge for single-answer datasets, which is also fully open-sourced to the community.
Data Studio: Core Crowdsourced Collaborative Data Ecosystem
AGI-Eval’s Data Studio serves as the backbone of its open community. The active collaborative data platform boasts over 30,000 crowdsourcing contributors, 485 categorized task labels and a total dataset volume exceeding 320,000 entries. It supports multiple data collection modes including single-sample annotation, data expansion and Arena competitive benchmark data, and enforces dual-layer machine + human auditing to guarantee strict data quality control. Users may browse and download public benchmark datasets or upload self-built evaluation corpora to co-build the open resource library. Private dataset hosting services are additionally provided for universities and research labs with advanced confidential benchmarking demands.
Frontier Multimodal Video Reasoning Benchmark: MMWorld Bench Co-Developed With UC Santa Cruz, UCSB & Microsoft Research
AGI-Eval partnered with the University of California, Santa Cruz, UC Santa Barbara and Microsoft Research to host the newly launched MMWorld Bench, a specialized benchmark designed to evaluate large multimodal models’ world simulation capabilities. The dataset spans 7 major sectors and 69 subfields including art & sports, commerce, science and medical healthcare, containing 1,910 high-quality video clips (average length 102 seconds) paired with 6,627 QA pairs. Unlike traditional video benchmarks limited to basic object recognition, MMWorld demands advanced high-order reasoning: phenomenon interpretation, counterfactual speculation, future state prediction, domain professional expertise and chronological logic comprehension.
Benchmark outcomes reveal clear bottlenecks for state-of-the-art multimodal models: even top-tier GPT-4o only achieved an overall accuracy of 62.54%. Its performance varies drastically across sectors — reaching 91.14% on commercial scenarios but dropping to merely 47.87% for art and sports, and 62.94% for embodied physical tasks. The findings expose core weaknesses of current multimodal systems in cross-disciplinary generalization and dynamic real-world comprehension.
Cutting-Edge Interactive Evaluation Research: AMemGym Solves Static Benchmark Reuse Bias
The team proposed an innovative interactive online evaluation framework AMemGym, the world’s first interactive strategic benchmark platform for dialogue agents, targeting the pervasive reuse bias flaw of traditional static testing. Experimental data showed a model configuration ranked 4th under static evaluation could rise to 1st in real interactive dialogue, while standard RAG systems often suffer ranking downgrades due to retrieval noise interference. AMemGym pioneered a three-stage diagnostic workflow: write — retrieve — utilize, enabling developers to precisely locate memory failure root causes: failed information storage, unsuccessful retrieval or erroneous information utilization. This research has been submitted to ICLR 2026, marking AGI-Eval’s research methodology at the international cutting edge of LLM evaluation.
Standardization & Globalization: CATArena AI Game Competition Co-Launched With SJTU & Meituan
To advance unified AI evaluation standards, AGI-Eval collaborated with Shanghai Jiao Tong University and Meituan to launch CATArena, an AI competitive arena built around strategic board games Gomoku and Texas Hold’em, which measures comprehensive general intelligence via iterative game competition. Benchmark results show domestic model Qwen 3 Coder ties GPT-5 for first place, while the Claude series renowned for universal reasoning failed to enter the top three. This proves CATArena evaluates far more than single-step logical deduction — it assesses end-to-end practical intelligence covering strategic coding, iterative self-learning and cross-scenario generalization.
Whether you are a developer conducting horizontal model capability comparisons, an enterprise team implementing model quality control, or an academic scholar researching LLM evaluation methodologies, this authoritative benchmark community jointly established by China’s top universities is an essential reference tool for all AI research and decision-making workflows.