17 Juillet

Kimi K3 is 2.8 trillion parameters and the open-weight gap just vanished

Moonshot AI dropped Kimi K3 on July 16. It is a 2.8-trillion-parameter sparse Mixture-of-Experts model with a 1-million-token context window, native vision, and an always-on reasoning mode. By every available metric it is the largest open-weight model ever released. It also competes head-to-head with the best closed models from Anthropic and OpenAI.

This was not supposed to happen yet. Analysts tracked by Fortune did not expect a Chinese lab to reach Fable 5 performance until early 2027. Moonshot delivered it five months early.

The timing is brutal. K3 landed hours before Google’s Gemini 3.5 Pro launch and the opening of the World AI Conference in Shanghai, where Xi Jinping delivered his first-ever WAIC keynote. Moonshot dropped a frontier-class model into every competitor’s press cycle and there is nothing accidental about the calendar.

The architecture

K3 is built on two new components Moonshot developed internally and published as open research on GitHub.

Kimi Delta Attention (KDA) is a hybrid linear attention mechanism that enables up to 6.3x faster decoding in million-token contexts. Attention Residuals (AttnRes) is a drop-in replacement for standard residual connections that delivers roughly 25 percent higher training efficiency at less than 2 percent additional cost.

The model ships in two variants. K3 Max handles chat and agent tasks. K3 Swarm Max targets large-scale parallel processing. Both are OpenAI SDK compatible, so switching is a base URL change. Context caching is automatic with no cache IDs, TTLs, or extra parameters to manage.

API pricing sits at $3 per million input tokens and $15 per million output. Cached input drops to $0.30 per million. That is not DeepSeek-cheap (V4 runs $0.87 per million output) but it annihilates Anthropic’s pricing on equivalent workloads. Fable 5 costs $50 per million output tokens. K3 offers 3x lower output pricing for near-equivalent performance.

Full model weights land on July 27, 2026.

The benchmarks

Artificial Analysis ran a private evaluation alongside public leaderboard data. The results are striking.

On GDPval-AA v2, which tests real-world tasks across 44 occupations and 9 industries, K3 scored 1,687. Third overall, behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), ahead of Claude Opus 4.8 (1,600).

On AA-Briefcase, a private agentic benchmark for long-horizon knowledge work, K3 took second with 1,527. It beat GPT-5.6 Sol Max (1,495) and trailed only Fable 5 Max (1,587).

BrowseComp, which measures long-horizon high-difficulty information seeking, gave K3 a state-of-the-art 91.2 out of 100.

K3 ranked first in 4 out of 8 real-world task automation benchmarks including Automation Bench, SpreadsheetBench 2, and BrowseComp. It took number one on Arena.AI’s Frontend Code Arena with 1,679, beating both Fable 5 and GPT-5.6 Sol.

Moonshot says it hit these numbers in a single-agent setup with no context compression or multi-agent workarounds. Just raw 1-million-token context and strong retrieval.

On coding benchmarks, K3 placed top three across six tests. It led SWE Marathon and Program Bench and trailed only GPT-5.6 Sol on Terminal Bench 2.1 by half a point.

Moonshot claimed K3 performs “competitively” with Anthropic Fable 5 and “substantially outperformed” Opus 4.8, GPT-5.6 Sol, and GPT-5.5. Those are company claims. The independent benchmark data backs them up well enough to take seriously.

The 48-hour chip demo

Beyond benchmarks, Moonshot showed a proof-of-concept that says more about their real ambition than any leaderboard score.

K3 was tasked with designing a physical chip to run a nano-scale version of itself. Over 48 hours of continuous autonomous operation, the model completed the full construction pipeline from architectural design through optimization and verification using open-source EDA tools. The result was a 4-square-millimeter chip design that hit timing convergence at 100 MHz and decoded over 8,700 tokens per second in simulation.

This is not a shipping product. It is a demonstration of sustained multi-day autonomous agency. Reading documentation, making design decisions, running verification loops, iterating on failures. That is a different category of capability from single-turn question answering.

In a second demo, K3 reproduced the universal I-Love-Q relation from computational astrophysics in roughly two hours. That calculation typically takes a senior researcher one to two weeks. K3 read and cross-validated over 20 papers and built a complete numerical pipeline.

Why this matters for pricing

The closed-model pricing premium just lost its last justification.

DeepSeek V4 stable lands July 24 at $0.44 per million output tokens. K3 weights drop July 27. By the end of this month, developers will have access to two frontier-class open models from Chinese labs, plus DeepSeek V4 Pro at 1.6T parameters, plus Qwen, GLM, and Thinking Machines’ Inkling.

Fable 5 at $50 per million output tokens requires a 70x price premium over DeepSeek for routine workloads. K3 does not close that gap. It eliminates the argument that the premium buys you meaningful capability separation. Every CFO running model evaluations is about to ask why they are paying it.

Cursor already uses Kimi models for Composer 2. DoorDash delegates lower-level coding work to Kimi K2.6. Thinking Machines used Kimi K2.5 to generate training data for Inkling. The model is already inside production stacks at Western companies.

The geopolitical dimension

Moonshot trained a 2.8-trillion-parameter model inside China’s compute ecosystem under US export controls. Huawei’s Atlas 950 SuperPoD is on the WAIC show floor this week. China is reportedly metering limited Nvidia H200 imports to fill compute gaps while domestic silicon scales.

The message from Shanghai this week is a full-stack pitch. Chips from Huawei. Models from Moonshot and DeepSeek. Governance rules from Xi’s proposed World AI Cooperation Organization. All packaged for global south delegations deciding whose AI stack to build on.

Anthropic has accused Moonshot, z.ai, Minimax, Alibaba, and DeepSeek of “illicit” distillation attacks. US politicians are drafting legislation to stop Chinese developers from distilling American models. Moonshot’s president, Yutong Zhang, framed the constraint differently at Davos this year. They knew they did not have the luxury to simply scale compute. That forced focus on fundamental research and efficiency. KDA and AttnRes are what that constraint produced.

Moonshot raised $2 billion in May at a valuation above $20 billion with over $200 million in annual recurring revenue. Backers include Alibaba, Tencent, Meituan, and HSG. An IPO in Hong Kong is reportedly in preparation. K3 is the pitch deck.

The open-source frontier gap is gone. What remains is a price war the closed labs cannot win on capability alone.

Mots-cles

kimi k3 moonshot ai open source llm china ai models mixture of experts