Research Methodology
Last updated: 2026-09-29
What this page covers
This page documents how China AI Hub turns scattered Chinese-language sources into the structured, English-language reference data on the rest of the site. It explains, in order: how data is collected, how it is verified, how claims are labeled, and how evidence is recorded. Editorial quality rules live on a separate page — the editorial standards — which this page links to and from.
1 · Data collection: official sources first
Every important factual claim must be traceable to a source, and sources are ranked by a four-tier hierarchy. A community-only source can never carry a critical fact (pricing, license, context window, or benchmark score) on its own.
| Tier | Type | Examples |
|---|---|---|
| L1 | Primary sources | Official docs, model cards, official pricing pages, papers, official GitHub, official announcements |
| L2 | High-quality technical | Academic papers, benchmark organizations, cloud-provider docs, independent technical research |
| L3 | Major industry media | Reuters, Bloomberg, Financial Times, TechCrunch, specialist publications |
| L4 | Community | GitHub discussions, Hugging Face, Reddit, developer communities |
L1 is required for pricing, API availability, license and official release dates. When a fact first surfaces in media, it is traced back to its L1 source before it is published.
2 · Verification workflow
Verification means an agent actually re-fetches the primary source and compares the extracted value against the stored value — it is not a guess. The per-update sequence is:
- Open the official source URL(s).
- Extract the exact statement or structured value.
- Compare it with the current stored value.
- If unchanged, refresh the
last_verifieddate. - If changed, create a change proposal recording old → new, effective date and source; pricing and benchmark changes are appended to history rather than silently overwritten.
- If the source is unreachable, leave the data unchanged and flag it for follow-up.
Each entity record carries a verification status: verified or
partially_verified. A record is partially_verified when a core field
(context window, architecture, parameters, capabilities, or benchmarks) is documented as not
publicly disclosed — the gap is stated, never hidden. Across the reference dataset this is
currently a 45 / 14 split: 45 records fully verified and 14
records partially_verified (13 models plus 1 API whose vendors decline to
disclose one or more core fields).
3 · Four claim labels
Every sentence that carries a fact or an interpretation is placed in one of four buckets, so the reader can always tell whose claim they are reading:
| Label | Meaning |
|---|---|
| Official fact | Stated in a primary (L1) source. |
| Vendor-reported claim | Published by the vendor (e.g. a benchmark score) and not independently re-verified. |
| Third-party evidence | From an independent evaluation or academic source. |
| China AI Hub analysis | Our own synthesis and interpretation, always introduced as analysis, never as fact. |
Analysis is kept out of the fact columns. Where a third-party label would be required but no third-party evidence exists, the page says so rather than implying one.
4 · The evidence layer: eight fields
Every source attached to an entity record is normalized into a single eight-field shape.
This is what the Data Hub (data.sinoaihub.com) exposes, and what the main site's
source tables are derived from.
| Field | Meaning | Rules |
|---|---|---|
evidence_id | Stable, globally-unique identifier | src-<collection>-<entity>-<n> |
source_name | Human-readable provenance label | — |
source_url | Canonical locator | — |
source_type | Type of source | See enumeration below |
published | Source publication date | Filled only where documented; — otherwise (never guessed) |
verified | Verification date | Formerly last_verified |
confidence | high | medium | low | — |
conflict | Contradiction flag | — unless the record documents a conflict between sources |
The source_type enumeration is: Official,
Official documentation, Model card, Vendor-reported,
Independent benchmark, China AI Hub analysis, and
Literature. Across the 81 entity records this produces 236 evidence rows.
5 · What is never published
Fabricated prices, benchmarks, parameters, context windows, capabilities, licenses, release dates, company details, customer counts or market share are never published. Neither are "best model" verdicts from a single benchmark, nor vendor-reported scores presented as independent findings. See the editorial standards for the full quality gate.