China AI Hub China AI Hub

Guides / Self-Hosting Chinese Open-Weight Models: Licenses, Hardware and Quantization

Self-Hosting Chinese Open-Weight Models: Licenses, Hardware and Quantization

Updated: 2026-09-29

Self-Hosting Chinese Open-Weight Models: Licenses, Hardware and Quantization
Image: AI-generated illustration (Seedream)

Short answer. Eleven of the twenty-one tracked models ship open weights, but “open weight” is not one legal category — the database records five license regimes, from DeepSeek’s unconditional MIT to MiniMax M2.7’s non-commercial-only terms with prior authorization required for any commercial use. Self-hosting is therefore a two-part decision: can you legally deploy the weights (license), and can you practically serve them (hardware + quantization). The least-friction route is a DeepSeek MIT or Zhipu Apache-2.0 checkpoint; the 1M-context open-weight option is currently Zhipu-only (GLM-5.3, GLM-5.3-Flash).

Decision criteria

CriteriaRelevance / Notes
License regimeMIT (DeepSeek) → Apache-2.0 (Zhipu) → conditional MIT-style with revenue/MAU triggers (Kimi K3, Qwen3.8-2.4T-A95B, MiniMax M3) → non-commercial-only (MiniMax M2.7)
Revenue / scale triggersKimi K3: MaaS >$20M/yr or >100M MAU; Qwen A95B: >100M MAU or >$50M/12mo MaaS; MiniMax M3: >$20M/yr
Total parameters1.6T (V4-Pro) to 320B (GLM-5.3-Flash) — total params dictate memory footprint even for sparse MoE
Active parameters8B–49B — the sparse-MoE inference cost driver
Context window (open)1M only on GLM-5.3 / GLM-5.3-Flash; Qwen A95B 256K; others below
Precision shippedFP8 (GLM-5.3, GLM-5.3-Flash), BF16/FP8 (GLM-5.2)
Quantization routeStandard route for single-GPU serving of 30B-class models; frontier MoE needs clusters
Serving ecosystemvLLM/Ollama via Qwen-Agent; DeepSeek Harness for DeepSeek models

The open-weight catalog

ModelLicenseParametersContextNotes
DeepSeek-V4-ProMIT1.6T / 49B active1MLargest open model; verify HF vs GA checkpoint
DeepSeek-V4.1-FlashMIT552B / 8–16B active1MSparse encoder-decoder; non-trivial to serve
DeepSeek-V3.2MIT61-layer DSA MoEnot statedDiscontinued; historical reference
GLM-5.3Apache-2.0744B / 40B active1MText-only; open 1M option
GLM-5.3-FlashApache-2.0320B / 18B active1MVision + video + computer use
GLM-5.2MIT744B / 40B active1M“Pure open, no regional limits”
Kimi K3Kimi K3 Licensenot disclosed1M / 1M outputConditional MIT-style
MiniMax-M3Community Licensenot disclosed1MAttribution + revenue trigger
MiniMax-M2.7Non-commercial onlynot disclosed200KStrictest terms
Qwen3.8-2.4T-A95BQwen3.8-Max License2.4T / 95B active256KWeights only, no API

Entity routing

Route by the license you can clear and the hardware you can field.

ScenarioBest-documented fitWhy
Least legal friction, largest open modelDeepSeek-V4-ProMIT, 1.6T/49B, 1M context
Cheapest open flash, 1M context + visionDeepSeek-V4.1-FlashMIT, 552B/8-16B active
Open-weight 1M context, text-onlyGLM-5.3Apache-2.0, 744B/40B
Open 1M multimodalGLM-5.3-Flash320B/18B, vision+video+computer-use
MIT “no regional limits” GLMGLM-5.2Deprecated but MIT; 744B/40B
Long-output 1M open modelKimi K3Conditional license, 1M/1M
Cheapest open flagshipMiniMax-M3Community license, attribution + trigger
Scale-triggered open model, no APIQwen3.8-2.4T-A95B256K, weights only
Historical MIT referenceDeepSeek-V3.2Discontinued, MIT

What the evidence shows

The license is a cost and a risk that appears on no pricing page, and it activates precisely when a business starts succeeding. The three conditional licenses share a structure: permissive everyday use with an obligation that scales with revenue, MAU or the Model-as-a-Service business model — differing only in where the line is drawn ($20M/yr MiniMax vs $20M/12mo Kimi MaaS vs $50M/12mo Qwen MaaS) and what is owed (attribution, notice, separate agreement, or prior authorization). MiniMax M2.7 is the strictest — non-commercial only, any commercial use requiring prior written authorization.

The thresholds are low enough to be real, not theoretical. A $20M annual revenue trigger is well within reach of a successful vertical SaaS product built on a Chinese open model. A company that hits the Kimi K3 MaaS trigger or the Qwen $50M/12-month MaaS trigger without having budgeted for a separate license negotiation has created an unplanned commercial dependency. The license, in other words, is part of total cost of ownership.

Hardware is the second filter. A 4-bit-quantized 30B-class open model runs on a single high-memory GPU; frontier-scale open models (744B–1.6T) need clusters even though MoE inference is sparse, because total parameters still dictate memory. Quantization is the standard route, and its quality is the adopter’s responsibility to evaluate. The serving ecosystem is thin but real: Qwen-Agent documents vLLM/Ollama connectivity, and DeepSeek Harness is the DeepSeek-native runtime.

China AI Hub analysis indicates the licensing landscape has shifted from a binary to a spectrum: the earlier generation of Chinese open models trended toward straightforward MIT or Apache, but the 2026 database shows the emergence of the conditional permissive license — MIT plus a monetization guardrail — which means “open weights” no longer implies “I can build a business on this without ever talking to the lab again.”

Selection procedure

Work through these steps in order.

  1. Clear the license first. DeepSeek MIT and Zhipu Apache-2.0 are the only regimes a large enterprise can clear without bespoke negotiation. If your revenue or MAU will cross a threshold ($20M/yr for MiniMax M3 and Kimi K3 MaaS, $50M/12mo for Qwen MaaS), the conditional licenses activate an obligation — budget for the conversation before you scale.
  2. Reject the non-commercial trap explicitly. MiniMax M2.7 is non-commercial only, with prior written authorization required for any commercial use. Do not build a product on it unless you have that authorization in hand.
  3. Size the hardware to total parameters, not active parameters. Even a sparse MoE must load its total parameters into memory: 1.6T (V4-Pro) is a cluster deployment, 320B (GLM-5.3-Flash) is large but more tractable. Quantization is the standard route, and its quality is yours to evaluate.
  4. If long context is non-negotiable and you must self-host, the field is Zhipu only. GLM-5.3 and GLM-5.3-Flash are the only open-weight 1M-context models; Alibaba’s open A95B is 256K.
  5. Pin the checkpoint before you depend on it. Verify the HF checkpoint matches the serving checkpoint (DeepSeek-V4-Pro’s was last modified 2026-06-22 vs an 0813 GA), and confirm Zhipu’s license applies to the actual weight artifacts, not just the repo.

Limitations

The license is quoted from vendor statements and repository metadata, not legal review — re-check the license file in the specific model repository before any commercial decision. Zhipu’s Apache-2.0 is read from GitHub repo metadata rather than a dedicated weights-license file, so a rigorous procurement must confirm it applies to the actual weight artifacts. The DeepSeek-V4-Pro HF checkpoint was last modified 2026-06-22 and it is not documented whether it matches the 0813 GA API checkpoint. Model-level capability flags trace to official pages; a missing flag is recorded as absence, not verified non-capability. Parameter counts are not disclosed for Moonshot and MiniMax models, so hardware planning for those is harder to pin down.

Sources

Labels used above: Official fact (license text, parameter counts, context windows and precision from primary sources), Vendor-reported claim (capability and serving statements published by the vendor), and China AI Hub analysis (our synthesis, introduced as such). No third-party evaluation evidence is currently recorded for these open-weight models.

Sources