Benchmarks / LiveCodeBench
LiveCodeBench
A contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, plus self-repair, code-execution and test-output-prediction tasks.
Short answer. LiveCodeBench is a contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, and grades models by executing their code against hidden tests. Its purpose is to give a coding score that cannot be gamed by memorising an old test set.
What it measures
LiveCodeBench measures real coding ability on problems that post-date a model’s training cutoff. It draws from three competitive-programming platforms and adds new problems over time, so the benchmark stays ahead of training-data contamination — the failure mode that makes older benchmarks like HumanEval and MBPP increasingly unreliable.
Beyond code generation, LiveCodeBench also measures a broader set of code capabilities: self-repair (fixing a failing program), code execution (predicting what a program outputs), and test-output prediction. This makes it a more holistic coding benchmark than a pure “write a function” suite.
How it works
LiveCodeBench collects problems from periodic contests on LeetCode, AtCoder and Codeforces, applying a time cutoff so that only problems published after a model’s knowledge cutoff are eligible. A model’s solution is then executed against hidden test cases, and a problem counts as solved only if the code passes the tests.
Scoring is pass@1 — the fraction of problems solved correctly on the first submission. Because pass@1 depends on sampling temperature and prompting, it is a single-pass signal and is not the same as pass@k or a “best-of-n” figure. The benchmark is hosted at livecodebench.github.io, with the problem set and toolkit in the LiveCodeBench/LiveCodeBench repository, and it has evaluated dozens of base and instruction-tuned models.
What it does NOT measure
LiveCodeBench does not measure general reasoning, knowledge or non-coding agentic skill. Pass@1 also has structural limits: it rewards a specific first-shot style and can understate models that improve with self-repair or longer sampling budgets. And because the problem set grows continuously, any score is tied to a specific snapshot — a “LiveCodeBench 60%” from one date is not the same measurement as a “LiveCodeBench 60%” from a year later.
Chinese model results
Not publicly documented. As of 2026-09-29, no model page in the China AI Hub database publishes a LiveCodeBench score. The tracked coding models — DeepSeek-V4.1-Flash, Kimi K2.7 Code, GLM-5.3-Flash and others — publish on Terminal-Bench, SWE-bench and DeepSWE instead. This page records no evaluations, and we do not import LiveCodeBench numbers from third-party leaderboards. Treat any LiveCodeBench figure seen elsewhere as outside this database until a tracked model publishes one.
How to read LiveCodeBench results
China AI Hub analysis: LiveCodeBench’s value is its contamination control, which makes it a more trustworthy coding signal than older static benchmarks — provided the score’s snapshot date and pass@1 setting are stated. For choosing a coding model, pair it with agentic-coding signals rather than reading it in isolation; see Choosing a coding model, Benchmark Methodology Divergence and Chinese AI coding models.
Contamination as a moving target
LiveCodeBench exists because the coding benchmark field had a specific failure mode: static suites like HumanEval and MBPP became saturated and memorisable, so their scores stopped meaning what they claimed. LiveCodeBench’s answer is structural — draw problems from recent contests after a cutoff date, so a model cannot have trained on them, and keep adding new problems as old ones age into training corpora. The benchmark is therefore not a fixed snapshot but a continuously refreshed one, which is both its strength and the reason a score must carry its snapshot date.
The same contamination logic applies across this database’s coding benchmarks: SWE-bench and DeepSWE test repository-scale and long-horizon software engineering, while LiveCodeBench tests competitive-programming problem solving under a pass@1 budget. They measure different things and should not be merged into a single “coding score”. See Choosing a coding model and Chinese AI coding models.
Related entities
- Models: DeepSeek-V4.1-Flash · Kimi K2.7 Code · GLM-5.3-Flash — the coding-focused models that would be the natural LiveCodeBench candidates.
- Comparison: DeepSeek-V4.1-Flash vs Qwen3.8-Flash.
- See all benchmarks.
Labels used on this page: Official fact (benchmark design and methodology from the primary paper and official repository), China AI Hub analysis (our synthesis of how to read scores, introduced as such). No vendor-reported claim or third-party evidence is recorded for this benchmark, because no tracked Chinese model has published a LiveCodeBench result.
Benchmark methodology
| Task type | Competitive-programming code generation plus self-repair, code execution and test-output prediction |
|---|---|
| Dataset size | Continuously growing; initially 300+ problems from LeetCode, AtCoder and Codeforces contests (from May 2023 onward, updated over time) |
| Evaluation method | Models solve fresh contest problems released after a cutoff date; graded by executing code against hidden test cases |
| Scoring | Pass@1 accuracy (% of problems solved correctly on the first submission) |
| Contamination notes | Problems are drawn from contests after a fixed cutoff date specifically to avoid contamination; the benchmark keeps adding new problems to stay ahead of training leakage. |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|
Limitations
Pass@1 depends on sampling temperature and prompting, and on the execution sandbox. Newer problems are added continuously, so a score is always tied to a specific snapshot. No Chinese model in the China AI Hub database currently publishes a LiveCodeBench score, so this page records no evaluations.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does LiveCodeBench measure?
A contamination-controlled coding benchmark that continuously collects fresh problems from LeetCode, AtCoder and Codeforces contests, plus self-repair, code-execution and test-output-prediction tasks.
Are LiveCodeBench scores independently verified?
Results on this page are labeled by source type (). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its LiveCodeBench data?
From 3 sources, last verified 2026-09-29.