Chinese AI Models 2026, Ranked

Home / Indexes / Technical Performance / Chinese AI Models 2026
Technical Performance · Report AI

Chinese AI Models 2026, Ranked

Two Chinese models now sit inside the global top four. But the release that matters most is the smallest one: on 14 August 2026 Alibaba shipped a 27.8-billion-parameter model, under Apache 2.0, that runs on a single consumer graphics card — and scores within seven points of the best model in the world. Every figure below is attributed, dated, and rated for confidence.

Thank you for reading this post, don't forget to subscribe!

By Report AI · Published 3 September 2026

52
Intelligence Index of a model that fits a 24GB GPU
~100×
smaller than Kimi K3, five points behind it
2
Chinese models in the global top four
3B
Alibaba model downloads — past Meta and Google
Analysis

The frontier stopped being the interesting part

The headline reading of 2026 is that China caught up. On the Artificial Analysis Intelligence Index, Alibaba’s Qwen3.8-Max scores 58 and Moonshot’s Kimi K3 scores 57.1, against 58.9 for OpenAI’s GPT-5.6 Sol. Two Chinese systems sit inside the global top four, and one of them ships its weights. That is a real change from eighteen months ago, and it is also the least interesting thing in the data.

The more consequential release came on 14 August 2026, and it is small. Qwen3.8-27B is a 27.8-billion-parameter dense model published under Apache 2.0. Quantised to four bits it occupies about 17GB — it runs on a single RTX 3090 or 4090, hardware that sells second-hand for a few hundred dollars. It scores 52. That is 6.9 points below the best model anyone has tested, from a model roughly 86× smaller than its own lab’s flagship and about 100× smaller than Kimi K3.

The competitive question has therefore moved. It is no longer “can Chinese labs reach the frontier” — they have, at least on composite benchmarks. It is “what is the frontier worth when 90% of its capability runs on a gaming GPU for free?” A closed model priced per token competes not only against other closed models but against a free artefact that a company can run inside its own firewall with no per-call cost and no vendor dependency at all. On agentic coding the gap is narrower still: Qwen3.8-27B posts 61.7 on SWE-bench Pro, against 64.7 for xAI’s Grok 4.5 — three points, between a downloadable 27B model and a frontier proprietary system.

Key takeaways
  • Parity at the top is real but narrow. Two Chinese models in the top four, both within two points of the leader — on one composite index.
  • The efficiency release matters more than the flagship. 52 points on a 24GB consumer card changes who can deploy frontier-adjacent AI, and at what cost.
  • Open weights are the strategic difference. Kimi K3 and Qwen3.8-27B ship downloadable weights; the US frontier does not. Alibaba’s models have passed 3 billion downloads, ahead of Meta and Google.

The ranking

Artificial Analysis Intelligence Index, a composite across reasoning, knowledge, mathematics and coding. Chinese models shaded. One index, one moment in time — read the caveats below before quoting any of it.

# Model Lab · Country Index Parameters Weights Confidence
1 Claude Fable 5 see caveat Anthropic · US 59.9 undisclosed Closed MEDIUM
2 GPT-5.6 Sol OpenAI · US 58.9 undisclosed Closed HIGH
3 Qwen3.8-Max Alibaba · China 58 2.4T total / 95B active (MoE) Closed MEDIUM
4 Kimi K3 Moonshot · China 57.1 2.8T (MoE) Open (MIT-based) HIGH
5 Claude Opus 4.8 Anthropic · US 55.7 undisclosed Closed HIGH
6 Qwen3.8-27B Alibaba · China 52 27.8B dense Apache 2.0 HIGH

Two caveats that matter more than the ordering.

Qwen3.8-Max’s score has moved three times. Artificial Analysis first published 53, then re-ran after reporting intermittent problems with the endpoint under test and got 56, and currently lists 58. We publish the current figure and the range, and rate it MEDIUM. A number that has moved five points in weeks is not a number to build an argument on.

Claude Fable 5 at the top requires explanation. Fable 5 was withdrawn within a week of launch in June 2026 under US export controls. Artificial Analysis lists the score with an Opus 4.8 fallback, which may mean the measurement partly reflects the fallback rather than Fable 5 itself. Until that is clarified we treat the top position as unsettled and rate it MEDIUM — a model you cannot use is a strange thing to call the best.

The Aug 14 release, in detail

Qwen3.8-27B is the model that changes the deployment maths. Its benchmark profile against models it has no business competing with:

Benchmark Qwen3.8-27B Compared with Confidence
Artificial Analysis Intelligence Index 52 GPT-5.6 Sol 58.9 · Kimi K3 57.1 HIGH
SWE-bench Pro (agentic coding) 61.7 Grok 4.5 64.7 · Claude Opus 4.8 69.2 MEDIUM
Terminal-Bench 2.1 73.0 Grok 4.5 83.3 MEDIUM
LiveCodeBench v6 90.3 MEDIUM
OSWorld-Verified 84.3 GPT-5.6 Sol 62.6 (OSWorld 2.0) MEDIUM

Specifications: 27.8B dense parameters, Apache 2.0, native context 262,144 tokens (extensible to ~1M via YaRN scaling), native image and video input. BF16 weights are ~55.6GB; at Q4_K_M quantisation it is 17.1GB, or 18.0GB with the vision tower loaded — inside a 24GB card at 64K context. MEDIUM

Read the OSWorld row carefully. The 84.3 and GPT-5.6’s 62.6 are different benchmark versions (OSWorld-Verified vs OSWorld 2.0) and are not directly comparable. We include both because the gap is the sort of thing that gets quoted out of context, and it should not be.

Capability per parameter

The same ranking, reordered by what it costs to get the score. This is where the Chinese open-weight strategy shows its shape.

Model Index Total params Index points per 100B params Runs on
Qwen3.8-27B 52 0.0278T 187.1 One 24GB consumer GPU
Qwen3.8-Max 58 2.4T 2.4 Datacentre / API only
Kimi K3 57.1 2.8T 2.0 Multi-GPU cluster
GPT-5.6 Sol · Claude Opus 4.8 58.9 · 55.7 undisclosed API only

The points-per-parameter column is DERIVED — index score divided by total parameters, computed by Report AI. Treat it as an illustration of the gap in deployment economics, not as a performance metric: dense and mixture-of-experts architectures are not comparable on total parameter count, since MoE models activate only a fraction per token (Qwen3.8-Max activates 95B of its 2.4T). Measured on active parameters the flagship gap narrows considerably. US models publish no parameter counts at all, so they cannot be placed in this column.

The rest of the Chinese field

Model Lab Known for Weights
DeepSeek V4-Pro DeepSeek Price-to-capability leader: MMLU-Pro 87.5, GPQA Diamond 90.1, SWE-bench Verified 80.6 at roughly $0.87 per million output tokens MIT
GLM-5 / 5.2 Zhipu AI Agentic and tool-use workflows; 77.8% SWE-bench Verified; reliable function calling and structured output Open
MiniMax MiniMax Very long context and mixed-media input — long documents, audio, video Open
Doubao ByteDance Not a benchmark contender but China’s consumer default, at 382M monthly users Closed

Sub-benchmark figures in this table are single-sourced and rated MEDIUM. Third-party leaderboards disagree materially about the ordering below the top four — one September 2026 board places MiniMax, GLM-5.3 and DeepSeek V4-Pro at ranks 16, 17 and 21 on a different composite entirely. We do not rank this group.

FAQ

Which Chinese AI model is best in 2026?

On the Artificial Analysis Intelligence Index, Alibaba’s Qwen3.8-Max at 58, just ahead of Moonshot’s Kimi K3 at 57.1. If you need downloadable weights, Kimi K3 is the strongest open model; if you need to self-host on modest hardware, Qwen3.8-27B.

What is the Alibaba model released on 14 August 2026?

Qwen3.8-27B — a 27.8-billion-parameter dense model under Apache 2.0, with 262K native context and image and video input. Quantised to four bits it is about 17GB and runs on a single 24GB consumer GPU, while scoring 52 on the Intelligence Index.

Can a 27B model really compete with frontier models?

On composite benchmarks it comes close — 52 against 58.9 for the leader. On agentic coding it is closer still: SWE-bench Pro 61.7 against 64.7 for Grok 4.5. It will not match frontier systems on the hardest reasoning, long-horizon agentic work, or reliability. What changed is the ratio: most of the capability, on hardware you own.

Have Chinese models overtaken US models?

Not on this index — the top two remain US systems. But the margin is under two points, and the strategic difference has moved elsewhere: China’s leading labs publish weights, and the US frontier labs do not.

Are these benchmark scores reliable?

Treat them as directional. Qwen3.8-Max’s index score moved 53 → 56 → 58 in weeks after endpoint problems and re-runs. Leaderboards disagree on ordering below the top four. Benchmark contamination is an open question across the field. We publish the numbers with their provenance so you can discount them yourself.

Methodology & sources

Index scores from Artificial Analysis model pages, read September 2026. Kimi K3 specifications and release dates from Moonshot AI’s announcement and coverage in VentureBeat, Axios and Tom’s Hardware. Qwen3.8-27B and Qwen3.8-Max specifications from Alibaba Cloud’s release materials and Artificial Analysis model pages; download milestone from Bloomberg, 15 August 2026. Comparison figures for GPT-5.6 Sol, Claude Opus 4.8 and Grok 4.5 from our own model ratings, July 2026. HIGH marks figures corroborated across two or more independent sources; MEDIUM marks single-sourced values, self-reported benchmarks, and any figure that has been revised; DERIVED marks the points-per-parameter column, which Report AI computes. Quantisation sizes vary by build and quantisation method — treat the 17.1GB figure as approximate. Corrections: see our methodology and corrections policy.