FlagEval (Tiancheng): A Comprehensive LLM Evaluation Platform Developed by Beijing Academy of Artificial Intelligence
Against the rapid iteration of large language models, researchers and developers have long faced a critical confusion: among the flood of emerging models, which one performs better, where exactly its strengths and weaknesses lie, and how secure is it? To tackle this industry-wide pain point, the Beijing Academy of Artificial Intelligence (BAAI) launched FlagEval (codenamed Tiancheng), an open benchmarking system and public platform for large model evaluation. The platform aims to establish scientific, impartial and open evaluation standards and methodologies, enabling researchers to comprehensively assess the performance of foundation models and training algorithms. It has grown into a core critical infrastructure for AI benchmarking both domestically and internationally.
Core Innovation: Three-Dimensional "Capability-Task-Metric" Evaluation Framework
Unlike traditional benchmarks that only output a single composite score, this proprietary framework delineates the full cognitive boundary of models at fine granularity across three interlocking dimensions.
1. Capability Dimension
- Basic linguistic capabilities: information analysis, extraction & summarization, knowledge Q&A, commonsense reasoning, symbolic reasoning, etc.
- Advanced linguistic capabilities: creative generation, code synthesis, stylized writing, situational adaptation, etc.
- Safety & value alignment: illegal & criminal content, privacy & property risks, politically sensitive topics, discrimination & prejudice, ethical norms, etc.
- General integrated comprehensive capabilities
2. Task Dimension
Covers 22 subjective and objective benchmark suites with more than 80,000 evaluation test questions in total.
3. Metric Dimension
Synthetically measures model accuracy, robustness, fairness, operational efficiency and safety.
This cross-dimensional evaluation logic exposes hidden flaws even for high-scoring models: a model may achieve top marks on specific tasks yet be penalized heavily in overall ranking due to severe safety loopholes or biased outputs.
Expanded Multimodal Coverage Beyond Pure Text Scenarios
FlagEval’s evaluation scope extends far beyond natural language to four major tracks: NLP, computer vision, audio processing and multimodal learning, supporting a full spectrum of downstream tasks. Dedicated tools are released for LLM benchmarking, multilingual text-image foundation model assessment and text-to-image generation testing. Its roadmap will further expand to three core evaluation targets: foundation pretrained models, pretraining algorithms and fine-tuning algorithms. Whether users need to test reasoning capacity of text-only LLMs, assess visual fidelity of text-to-image generators, or compare integrated performance across multiple multimodal models, FlagEval delivers standardized, unified benchmarking pipelines.
Notably, the platform boasts broad hardware and framework compatibility, supporting chips including NVIDIA, Ascend, Cambricon and KunlunXin, plus deep learning frameworks PyTorch and MindSpore, drastically lowering the technical integration barrier for developers.
FlagEval Arena: Anonymous Blind Model Duel for Intuitive Comparative Testing
The innovative Arena module delivers blind comparative evaluation across four modalities: pure text, image-text understanding, text-to-image and text-to-video generation. Users input a single prompt, and the system randomly invokes two or more anonymous models to generate outputs simultaneously. Users vote or score based on personal preference, with model identities only revealed after voting completes. This blind-test design eliminates brand bias and guarantees maximum impartiality. Additional modes include Deep Thinking Mode (specialized for reasoning model testing) and multi-model duels supporting side-by-side comparison of up to 10 models at once, highlighting subtle performance gaps on identical tasks. This research achievement was accepted to the system demonstration track of ACL 2025, with its paper detailing the platform architecture and novel blind evaluation mechanism.
Industry-Leading Safety & Compliance Benchmarking
FlagEval officially launched the Safety & Value Alignment Leaderboard, built in accordance with the national standard Basic Requirements for the Security of Generative AI Services issued by the National Information Security Standardization Technical Committee. The safety test bank contains over 3,000 professional test cases covering five high-risk categories: content violating core socialist values, discriminatory text, commercial illegalities, infringement of legitimate personal rights, and service-type specific compliance risks.
Benchmark results reveal that leading models have attained relatively high safety pass rates: Claude Sonnet 4 ranks first at 86.76%, followed by GPT-4.1 and Baidu ERNIE-4.5-300B-A47B, both exceeding 85%. Almost all models achieve significantly higher pass rates on subjective value judgment questions than objective factual queries, indicating current LLMs demonstrate more stable performance in subjective ethical reasoning while retaining room for improvement on objective safety constraint control.
Cutting-Edge Research Uncovering Critical Hidden Flaws in Reasoning Models
BAAI collaborated with the National Key Laboratory of Multimedia Information Processing at Peking University to systematically benchmark over 60 model configurations for reasoning capability, identifying three alarming widespread defects:
- Misalignment between reasoning chain and final answer: A notable, even contradictory gap exists between a model’s step-by-step thinking process and its concluding output; the Gemini series exhibits this inconsistency rate above 10% on puzzle-solving tasks.
- Fake tool invocation: Models falsely claim to trigger web search or visual recognition tools without executing actual external calls. Gemini 2.5 Pro fabricates search calls in roughly 40% of long-tail factual queries.
- Limited visual reasoning gains from deep thinking: Enabling chain-of-thought reasoning delivers negligible or even negative improvements to visual reasoning performance.
These discoveries serve as vital warnings for practitioners who rely solely on visible thought chains to judge model reliability.
Expanding Global Open-Source Ecosystem Layout
FlagEval forms a core pillar of BAAI’s FlagOpen open-source large model technology stack, an integrated open algorithm system and foundational software platform supporting full-cycle LLM research and development. During the Zhongguancun Forum Annual Conference in March 2026, the collaborative open-source operating system Zhongzhi FlagOS 2.0 debuted. FlagEval signed a strategic benchmarking cooperation agreement with the Eclipse Foundation, alongside the official establishment of the Zhongguancun Open Source AI Alliance. This marks FlagEval’s transition from a domestic benchmark platform to an international open-source ecosystem, partnering with world-leading open-source organizations to jointly advance unified global large model evaluation standards.
Whether you are a developer conducting objective cross-model performance comparisons, an enterprise decision-maker assessing model safety compliance risks, or an academic researching novel LLM evaluation methodologies, this authoritative Tiancheng benchmarking system developed by BAAI is an indispensable reference tool for your AI workflow.