China AI Hub China AI Hub

Benchmarks / AIME

AIME

The American Invitational Mathematics Examination, repurposed as a benchmark: 15 integer-answer math problems used to test a model's advanced mathematical reasoning, with scores typically reported for AIME 2024 or AIME 2025.

Short answer. AIME is the American Invitational Mathematics Examination — a 15-problem, integer-answer competition exam — repurposed as a benchmark of advanced mathematical reasoning. Because it is difficult, exact-answer and widely reported by frontier model builders (typically as AIME 2024 or AIME 2025), it became a de-facto standard for “how good is this model at hard math”.

What it measures

AIME measures advanced competition mathematics: algebra, number theory, combinatorics and geometry at a level well above standard benchmarks. Its problems are not multiple-choice — each requires a derived integer answer, which means a model must actually reason to a correct result rather than rank four options. This makes AIME a stronger signal of genuine mathematical ability than a multiple-choice math test.

As a benchmark, AIME is usually reported as accuracy on a specific year’s exam. The 2024 and 2025 papers are the most common targets in modern model cards, so “AIME 2024” and “AIME 2025” function as two distinct, year-labelled sub-benchmarks.

How it works

Each AIME exam has 15 problems with integer answers from 000 to 999. Grading is exact-match: a model’s answer must equal the reference integer to count. There is no partial credit, and the raw score runs from 0 to 15; benchmarks typically report this as a percentage or as “X/15 correct”.

Because AIME is a fixed public exam administered by the Mathematical Association of America (MAA), the benchmark has no versioning of its own — the year is the version. This is the single most important fact when reading AIME numbers: an “AIME 90%” that does not state the year is uninterpretable.

What it does NOT measure

AIME does not measure general language ability, coding, agentic skill or breadth of knowledge — it is a narrow, deep test of competition math. Its exact-match grading gives no partial credit, so a model that reasons correctly but makes one arithmetic slip scores the same as one that cannot start. And because the exam is public, a model trained after a given year may have memorised that year’s problems, which is why the freshest exam (AIME 2025) is the more contamination-resistant target.

Chinese model results

Not publicly documented. As of 2026-09-29, no model page in the China AI Hub database publishes an AIME score. The tracked reasoning models — DeepSeek-V4-Pro, Kimi K3, Qwen3.8-Max and others — publish on GPQA Diamond, HLE and coding benchmarks, but not on AIME. This page records no evaluations, and we do not import AIME numbers from third-party leaderboards or vendor reports. Treat any AIME figure seen elsewhere as outside this database until a tracked model publishes one.

How to read AIME results

China AI Hub analysis: AIME is a high-signal but narrow math metric, and its scores are only meaningful when the year is stated and matched against a model’s training cutoff. It is best used to compare mathematical-reasoning strength, not overall model quality. The same year-and-mode discipline applies here as elsewhere in the database: see Benchmark Methodology Divergence and How to read vendor-reported benchmarks.

AIME as a year-labelled benchmark

Because AIME is a public exam administered annually (as AIME I and AIME II) rather than a versioned benchmark, the year is the version. In practice this matters more than it sounds: as frontier reasoning models began scoring at or near the ceiling on the older papers, evaluators shifted toward the freshest exam — AIME 2025 — as the harder, less-contaminated target. A score reported as simply “AIME” without a year therefore conflates measurements of very different difficulty.

The same dynamic explains why AIME became a reasoning benchmark rather than a knowledge benchmark. Its integer-answer, no-partial-credit format rewards models that can derive a correct answer end-to-end, which is exactly what reasoning models were built for. That is also its limit: AIME says little about a model’s breadth, coding, or real-world usefulness, only about its competition-math reasoning depth. For reading math-related claims elsewhere in this database, the same version-discipline applies — see Benchmark Methodology Divergence.

Labels used on this page: Official fact (exam format and grading from the MAA), China AI Hub analysis (our synthesis of how to read year-labelled scores, introduced as such). No vendor-reported claim or third-party evidence is recorded for this benchmark, because no tracked Chinese model has published an AIME result.

Benchmark methodology

Methodology of AIME
Task type Advanced competition mathematics — 15 integer-answer problems
Dataset size 15 problems per exam (integer answers from 000 to 999); benchmark scores typically report AIME 2024 or AIME 2025
Evaluation method Models answer each problem; graded by exact match against the integer answer
Scoring Accuracy (% of 15 problems answered correctly); the raw exam score is 0–15
Contamination notes AIME is a fixed public exam, so models trained after a given year may have seen its problems; vendors differ in which year(s) they report.
Last verified: · Data status: Current · Next review:

Results

Results for AIME
Model Model version Score Metric Date Source type Source

Limitations

AIME is an exam, not a versioned benchmark, so scores are only comparable when the same year is reported. Exact-match integer grading gives no partial credit for correct reasoning with a wrong final answer. No Chinese model in the China AI Hub database currently publishes an AIME score, so this page records no evaluations.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does AIME measure?

The American Invitational Mathematics Examination, repurposed as a benchmark: 15 integer-answer math problems used to test a model's advanced mathematical reasoning, with scores typically reported for AIME 2024 or AIME 2025.

Are AIME scores independently verified?

Results on this page are labeled by source type (). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its AIME data?

From 2 sources, last verified 2026-09-29.