China AI Hub China AI Hub

Benchmarks / LiveCodeBench

LiveCodeBench

A contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, plus self-repair, code-execution and test-output-prediction tasks.

Short answer. LiveCodeBench is a contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, and grades models by executing their code against hidden tests. Its purpose is to give a coding score that cannot be gamed by memorising an old test set.

What it measures

LiveCodeBench measures real coding ability on problems that post-date a model’s training cutoff. It draws from three competitive-programming platforms and adds new problems over time, so the benchmark stays ahead of training-data contamination — the failure mode that makes older benchmarks like HumanEval and MBPP increasingly unreliable.

Beyond code generation, LiveCodeBench also measures a broader set of code capabilities: self-repair (fixing a failing program), code execution (predicting what a program outputs), and test-output prediction. This makes it a more holistic coding benchmark than a pure “write a function” suite.

How it works

LiveCodeBench collects problems from periodic contests on LeetCode, AtCoder and Codeforces, applying a time cutoff so that only problems published after a model’s knowledge cutoff are eligible. A model’s solution is then executed against hidden test cases, and a problem counts as solved only if the code passes the tests.

Scoring is pass@1 — the fraction of problems solved correctly on the first submission. Because pass@1 depends on sampling temperature and prompting, it is a single-pass signal and is not the same as pass@k or a “best-of-n” figure. The benchmark is hosted at livecodebench.github.io, with the problem set and toolkit in the LiveCodeBench/LiveCodeBench repository, and it has evaluated dozens of base and instruction-tuned models.

What it does NOT measure

LiveCodeBench does not measure general reasoning, knowledge or non-coding agentic skill. Pass@1 also has structural limits: it rewards a specific first-shot style and can understate models that improve with self-repair or longer sampling budgets. And because the problem set grows continuously, any score is tied to a specific snapshot — a “LiveCodeBench 60%” from one date is not the same measurement as a “LiveCodeBench 60%” from a year later.

Chinese model results

Not publicly documented. As of 2026-09-29, no model page in the China AI Hub database publishes a LiveCodeBench score. The tracked coding models — DeepSeek-V4.1-Flash, Kimi K2.7 Code, GLM-5.3-Flash and others — publish on Terminal-Bench, SWE-bench and DeepSWE instead. This page records no evaluations, and we do not import LiveCodeBench numbers from third-party leaderboards. Treat any LiveCodeBench figure seen elsewhere as outside this database until a tracked model publishes one.

How to read LiveCodeBench results

China AI Hub analysis: LiveCodeBench’s value is its contamination control, which makes it a more trustworthy coding signal than older static benchmarks — provided the score’s snapshot date and pass@1 setting are stated. For choosing a coding model, pair it with agentic-coding signals rather than reading it in isolation; see Choosing a coding model, Benchmark Methodology Divergence and Chinese AI coding models.

Contamination as a moving target

LiveCodeBench exists because the coding benchmark field had a specific failure mode: static suites like HumanEval and MBPP became saturated and memorisable, so their scores stopped meaning what they claimed. LiveCodeBench’s answer is structural — draw problems from recent contests after a cutoff date, so a model cannot have trained on them, and keep adding new problems as old ones age into training corpora. The benchmark is therefore not a fixed snapshot but a continuously refreshed one, which is both its strength and the reason a score must carry its snapshot date.

The same contamination logic applies across this database’s coding benchmarks: SWE-bench and DeepSWE test repository-scale and long-horizon software engineering, while LiveCodeBench tests competitive-programming problem solving under a pass@1 budget. They measure different things and should not be merged into a single “coding score”. See Choosing a coding model and Chinese AI coding models.

Labels used on this page: Official fact (benchmark design and methodology from the primary paper and official repository), China AI Hub analysis (our synthesis of how to read scores, introduced as such). No vendor-reported claim or third-party evidence is recorded for this benchmark, because no tracked Chinese model has published a LiveCodeBench result.

Benchmark methodology

Methodology of LiveCodeBench
Task type Competitive-programming code generation plus self-repair, code execution and test-output prediction
Dataset size Continuously growing; initially 300+ problems from LeetCode, AtCoder and Codeforces contests (from May 2023 onward, updated over time)
Evaluation method Models solve fresh contest problems released after a cutoff date; graded by executing code against hidden test cases
Scoring Pass@1 accuracy (% of problems solved correctly on the first submission)
Contamination notes Problems are drawn from contests after a fixed cutoff date specifically to avoid contamination; the benchmark keeps adding new problems to stay ahead of training leakage.
Last verified: · Data status: Current · Next review:

Results

Results for LiveCodeBench
Model Model version Score Metric Date Source type Source

Limitations

Pass@1 depends on sampling temperature and prompting, and on the execution sandbox. Newer problems are added continuously, so a score is always tied to a specific snapshot. No Chinese model in the China AI Hub database currently publishes a LiveCodeBench score, so this page records no evaluations.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does LiveCodeBench measure?

A contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, plus self-repair, code-execution and test-output-prediction tasks.

Are LiveCodeBench scores independently verified?

Results on this page are labeled by source type (). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its LiveCodeBench data?

From 3 sources, last verified 2026-09-29.