China AI Hub China AI Hub

Benchmarks / ARC-AGI

ARC-AGI

The Abstraction and Reasoning Corpus for AGI: grid-based visual reasoning tasks that are trivial for humans but have historically defeated frontier language models, used to measure general fluid intelligence.

Short answer. ARC-AGI is the Abstraction and Reasoning Corpus, a benchmark of grid-based visual reasoning tasks that are trivially easy for humans but have historically defeated even frontier language models. It is designed to measure general fluid intelligence — learning a rule from a few examples and applying it to a new case — rather than memorised knowledge.

What it measures

ARC-AGI measures abstraction and few-shot generalisation. Each task shows a small set of input grids and their corresponding output grids, and the model must infer the underlying transformation and produce the correct output grid for unseen test inputs. The transformations are things humans grasp instantly — pattern completion, symmetry, counting, object manipulation — but they are deliberately built on priors that are “as close as possible to innate human priors”, in François Chollet’s framing.

The benchmark’s central claim is that scaling and memorisation do not by themselves produce this kind of fluid intelligence, which is why ARC-AGI became a reference for the “is scaling enough to reach AGI?” debate.

How it works

A task is counted as solved only if the model produces the correct output grid for every test input in the task, including the correct grid dimensions. Each test input allows up to 3 trials — a standardised rule that applies to both humans and AI systems.

Scoring is accuracy: the percentage of tasks solved. The original release (ARC-AGI-1) contains 400 training tasks and 400 public evaluation tasks. Newer families — ARC-AGI-2 and ARC-AGI-3 — introduce fresh task sets specifically to counter overfitting to ARC-AGI-1, so scores are only comparable within a single benchmark version. The benchmark is maintained in the fchollet/ARC-AGI repository and by ARC Prize (arcprize.org).

What it does NOT measure

ARC-AGI is a deliberately narrow probe of a specific cognitive ability. It does not measure language, knowledge, coding, or useful real-world task performance, and success on ARC-AGI does not translate directly into practical assistant quality. Conversely, a poor ARC-AGI score does not mean a model is useless — it means the model lacks a particular kind of fluid abstraction that humans have and most LLMs historically lacked.

Chinese model results

Not publicly documented. As of 2026-09-29, no model page in the China AI Hub database publishes an ARC-AGI score. The tracked reasoning models — DeepSeek-V4-Pro, Qwen3.8-Max, Kimi K3 and others — publish on GPQA Diamond, HLE and coding benchmarks, but not on ARC-AGI. This page records no evaluations, and we do not import ARC-AGI numbers from third-party sources. Treat any ARC-AGI figure seen elsewhere as outside this database until a tracked model publishes one.

How to read ARC-AGI results

China AI Hub analysis: ARC-AGI is best read as a measure of fluid abstraction, not overall model quality, and only within a single benchmark version. Its value is diagnostic — separating memorisation-driven performance from generalisation — rather than as a buying criterion. See Benchmark Methodology Divergence, How to read vendor-reported benchmarks and the state of China’s AI models.

ARC-AGI and the AGI debate

ARC-AGI is tied to a specific theoretical position. Its author, François Chollet, argues in On the Measure of Intelligence that intelligence is skill-acquisition efficiency — the ability to learn a new rule from few examples — rather than accumulated skill, and that scaling and memorisation do not by themselves produce it. ARC-AGI is the operational test of that claim: tasks whose rules a human grasps almost instantly but which defeated large language models for years.

That framing is why ARC-AGI’s score should be read differently from a knowledge or coding benchmark. A model that scores near zero on ARC-AGI-1 can still be an excellent assistant, and a model that scores well demonstrates a specific kind of fluid abstraction, not general competence. As reasoning systems began to close the gap on ARC-AGI-1, the benchmark released harder successor families — ARC-AGI-2 and ARC-AGI-3 — which is the standard arms race between a benchmark and the models it measures, and another reason scores are only meaningful within one version. See ARC Prize for the current leaderboard.

Labels used on this page: Official fact (benchmark design, scoring and task format from the primary paper and official repository), China AI Hub analysis (our synthesis of what ARC-AGI does and does not tell you, introduced as such). No vendor-reported claim or third-party evidence is recorded for this benchmark, because no tracked Chinese model has published an ARC-AGI result.

Benchmark methodology

Methodology of ARC-AGI
Task type Grid-based abstract visual reasoning — infer an input→output transformation and apply it to new grids
Dataset size ARC-AGI-1: 400 training + 400 public evaluation tasks; newer families (ARC-AGI-2, ARC-AGI-3) add fresh task sets
Evaluation method A task is solved if the model produces the correct output grid for every test input (up to 3 trials per input)
Scoring Accuracy (% of tasks solved); a task counts only if all test inputs are correct
Contamination notes Evaluation tasks are held out; ARC-AGI-2 and ARC-AGI-3 introduce new task families to counter overfitting to ARC-AGI-1.
Last verified: · Data status: Current · Next review:

Results

Results for ARC-AGI
Model Model version Score Metric Date Source type Source

Limitations

ARC-AGI measures a specific form of abstraction, not general model quality; visual-grid reasoning is only weakly correlated with useful real-world task performance. No Chinese model in the China AI Hub database currently publishes an ARC-AGI score, so this page records no evaluations.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does ARC-AGI measure?

The Abstraction and Reasoning Corpus for AGI: grid-based visual reasoning tasks that are trivial for humans but have historically defeated frontier language models, used to measure general fluid intelligence.

Are ARC-AGI scores independently verified?

Results on this page are labeled by source type (). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its ARC-AGI data?

From 3 sources, last verified 2026-09-29.