China AI Hub China AI Hub

Research Methodology

Last updated: 2026-09-29

What this page covers

This page documents how China AI Hub turns scattered Chinese-language sources into the structured, English-language reference data on the rest of the site. It explains, in order: how data is collected, how it is verified, how claims are labeled, and how evidence is recorded. Editorial quality rules live on a separate page — the editorial standards — which this page links to and from.

1 · Data collection: official sources first

Every important factual claim must be traceable to a source, and sources are ranked by a four-tier hierarchy. A community-only source can never carry a critical fact (pricing, license, context window, or benchmark score) on its own.

Source tier hierarchy
TierTypeExamples
L1Primary sourcesOfficial docs, model cards, official pricing pages, papers, official GitHub, official announcements
L2High-quality technicalAcademic papers, benchmark organizations, cloud-provider docs, independent technical research
L3Major industry mediaReuters, Bloomberg, Financial Times, TechCrunch, specialist publications
L4CommunityGitHub discussions, Hugging Face, Reddit, developer communities

L1 is required for pricing, API availability, license and official release dates. When a fact first surfaces in media, it is traced back to its L1 source before it is published.

2 · Verification workflow

Verification means an agent actually re-fetches the primary source and compares the extracted value against the stored value — it is not a guess. The per-update sequence is:

  1. Open the official source URL(s).
  2. Extract the exact statement or structured value.
  3. Compare it with the current stored value.
  4. If unchanged, refresh the last_verified date.
  5. If changed, create a change proposal recording old → new, effective date and source; pricing and benchmark changes are appended to history rather than silently overwritten.
  6. If the source is unreachable, leave the data unchanged and flag it for follow-up.

Each entity record carries a verification status: verified or partially_verified. A record is partially_verified when a core field (context window, architecture, parameters, capabilities, or benchmarks) is documented as not publicly disclosed — the gap is stated, never hidden. Across the reference dataset this is currently a 45 / 14 split: 45 records fully verified and 14 records partially_verified (13 models plus 1 API whose vendors decline to disclose one or more core fields).

3 · Four claim labels

Every sentence that carries a fact or an interpretation is placed in one of four buckets, so the reader can always tell whose claim they are reading:

The four claim labels
LabelMeaning
Official factStated in a primary (L1) source.
Vendor-reported claimPublished by the vendor (e.g. a benchmark score) and not independently re-verified.
Third-party evidenceFrom an independent evaluation or academic source.
China AI Hub analysisOur own synthesis and interpretation, always introduced as analysis, never as fact.

Analysis is kept out of the fact columns. Where a third-party label would be required but no third-party evidence exists, the page says so rather than implying one.

4 · The evidence layer: eight fields

Every source attached to an entity record is normalized into a single eight-field shape. This is what the Data Hub (data.sinoaihub.com) exposes, and what the main site's source tables are derived from.

The eight evidence fields
FieldMeaningRules
evidence_idStable, globally-unique identifiersrc-<collection>-<entity>-<n>
source_nameHuman-readable provenance label—
source_urlCanonical locator—
source_typeType of sourceSee enumeration below
publishedSource publication dateFilled only where documented; — otherwise (never guessed)
verifiedVerification dateFormerly last_verified
confidencehigh | medium | low—
conflictContradiction flag— unless the record documents a conflict between sources

The source_type enumeration is: Official, Official documentation, Model card, Vendor-reported, Independent benchmark, China AI Hub analysis, and Literature. Across the 81 entity records this produces 236 evidence rows.

5 · What is never published

Fabricated prices, benchmarks, parameters, context windows, capabilities, licenses, release dates, company details, customer counts or market share are never published. Neither are "best model" verdicts from a single benchmark, nor vendor-reported scores presented as independent findings. See the editorial standards for the full quality gate.