China AI Hub China AI Hub

Benchmarks / SuperCLUE

SuperCLUE

A Chinese comprehensive evaluation system for large models, organised around four ability quadrants — language understanding and generation, professional knowledge, agent capability and safety — refined into 12 base capabilities.

Short answer. SuperCLUE is a Chinese comprehensive evaluation system for large language models, organised around four ability quadrants — language understanding and generation, professional knowledge and skills, agent capability, and safety — refined into 12 base capabilities. It is the CLUE benchmark team’s flagship, and its distinguishing feature is that the CLUE team runs the evaluation itself, including multi-turn open questions and agent tasks.

What it measures

SuperCLUE measures a Chinese model’s overall capability rather than a single skill. Its four quadrants cover: (1) language understanding and generation; (2) professional skills and knowledge; (3) agent capability — tool use and task planning; and (4) safety. These expand into 12 base capabilities, and the benchmark has also produced specialised tracks such as SuperCLUE-Agent (Chinese-native agent tasks) and SuperCLUE-Safety (multi-turn adversarial safety).

The result is a benchmark that tries to answer “how good is this Chinese model, all things considered” — a different question from the single-skill benchmarks like C-Eval (knowledge) or a coding benchmark. That breadth is its value, and also the reason its headline number needs careful reading.

How it works

SuperCLUE is a multi-part, team-run evaluation. Unlike a static multiple-choice dataset, it combines objective questions with multi-turn open-ended questions and agent/tool-use tasks, and the CLUE team scores the responses rather than asking vendors to self-report. The headline output is a composite total score (总分), with sub-scores for components such as OPEN multi-turn, open-ended questions and objective questions.

Because it runs on a periodic leaderboard cycle, SuperCLUE scores are time-stamped snapshots: a September leaderboard and a December leaderboard are different evaluations and are not directly comparable. The official website (superclueai.com) and the CLUEbenchmark/SuperCLUE repository publish the methodology and the rankings.

What it does NOT measure

Three limits are worth stating. First, the composite total hides capability-specific strengths — a model strong in professional knowledge but weak in agent tool-use can have the same total as one with the opposite profile. Second, because the CLUE team’s rubric and human scoring evolve between releases, cross-version comparisons are unreliable. Third, SuperCLUE is a Chinese-language evaluation: it does not measure English performance, coding depth, or the long-context and reasoning regimes that international benchmarks target.

Chinese model results

Not publicly documented. As of 2026-09-29, no model page in the China AI Hub database publishes a SuperCLUE score. The tracked models — Doubao Seed 2.1 Pro, GLM-5.3, Qwen3.8-Max and others — publish on international and coding benchmarks, not on SuperCLUE. This page therefore records no evaluations, and we do not import SuperCLUE numbers from the public leaderboard. Treat any SuperCLUE figure seen elsewhere as outside this database until a tracked model publishes one.

How to read SuperCLUE results

China AI Hub analysis: SuperCLUE’s composite is best read by its quadrant sub-scores, not the headline total, and only within a single leaderboard snapshot. Its agent and safety tracks are its most distinctive contributions, because they measure capabilities that pure multiple-choice benchmarks miss. As with every benchmark here, confirm the version and date before comparing; see Benchmark Methodology Divergence and How to read vendor-reported benchmarks.

SuperCLUE’s distinctive role

SuperCLUE is the third pillar of Chinese-native evaluation, and it plays a role the two static datasets do not. Where C-Eval and CMMLU are fixed multiple-choice corpora whose scores are self-reported, SuperCLUE is a team-run, periodically released evaluation that mixes objective questions with multi-turn open questions and agent/tool-use tasks — and the CLUE team grades the answers itself. That makes it closer in spirit to a human-graded arena than to a downloadable dataset, with the caveat that its rubric and human scoring evolve between releases.

Its most distinctive contributions are the agent and safety tracks (SuperCLUE-Agent and SuperCLUE-Safety), which measure capabilities — tool use, task planning, multi-turn adversarial robustness — that a pure knowledge benchmark cannot see. This is why SuperCLUE’s headline composite is best read through its quadrant sub-scores: two models with the same total can be strong and weak in opposite quadrants, and only the breakdown reveals that.

Labels used on this page: Official fact (benchmark design and methodology from the primary paper and official repository), China AI Hub analysis (our synthesis of how to read the composite score, introduced as such). No vendor-reported claim or third-party evidence is recorded for this benchmark, because no tracked Chinese model has published a SuperCLUE result.

Benchmark methodology

Methodology of SuperCLUE
Task type Chinese multi-part evaluation spanning language understanding/generation, professional knowledge, agent capability and safety
Dataset size Multi-part rolling evaluation across four ability quadrants and 12 base capabilities; no fixed single question count (periodic leaderboard releases)
Evaluation method Combines objective multiple-choice questions with multi-turn open-ended questions and agent/tool-use tasks; scored by the CLUE team rather than self-reported
Scoring Composite total score (总分) with sub-scores per quadrant (e.g. OPEN multi-turn, open questions, objective questions)
Contamination notes SuperCLUE maintains public and hidden test sets and runs its own evaluation, reducing reliance on vendor self-reported scores.
Last verified: · Data status: Current · Next review:

Results

Results for SuperCLUE
Model Model version Score Metric Date Source type Source

Limitations

The composite score aggregates heterogeneous sub-tasks, so a single total hides capability-specific strengths. Leaderboard snapshots are time-stamped and not comparable across dates. No Chinese model in the China AI Hub database currently publishes a SuperCLUE score, so this page records no evaluations.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does SuperCLUE measure?

A Chinese comprehensive evaluation system for large models, organised around four ability quadrants — language understanding and generation, professional knowledge, agent capability and safety — refined into 12 base capabilities.

Are SuperCLUE scores independently verified?

Results on this page are labeled by source type (). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its SuperCLUE data?

From 3 sources, last verified 2026-09-29.