Moonshot AI's K3 matches GPT-5.5 and Claude Opus 4.8 on independent benchmarks, but its pricing signals the end of ultra-cheap Chinese frontier AI.

Moonshot AI has released Kimi K3, a 2.8-trillion-parameter multimodal model that independent testing places on par with OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.8 — while pricing it well above its own predecessor and other Chinese open-weight rivals [1].

The launch matters for developers and AI buyers because K3 is the first open model in the roughly 3-trillion-parameter range, according to Kimi, and its full weights are scheduled for public release by July 27 [1]. At the same time, its token prices represent a sharp break from the rock-bottom rates that have defined Chinese AI offerings over the past year [1].

What the benchmarks show

Independent testing lab Artificial Analysis scored K3 at 57 on its Intelligence Index, putting it level with Opus 4.8 and GPT-5.5 but behind Claude Fable 5 and GPT-5.6 Sol [1]. That result broadly matches Kimi’s own benchmark claims, though Kimi’s internal tests were run at maximum or high thinking intensity and used different agent systems — KimiCode, Claude Code, or Codex — depending on the task, meaning conditions were not identical across all 35 tests [1].

On agentic tasks, Artificial Analysis gives K3 an Elo rating of 1,668 on GDPval v2, up from 1,190 for its predecessor K2.6 [1]. That puts it ahead of GLM-5.2 (1,514), GPT-5.5 (1,494), and Claude Opus 4.8 (1,600), though Claude Fable 5 still leads at 1,760 [1]. K3 also tops AutomationBench-AA, Artificial Analysis’s agentic software-as-a-service workflow evaluation, with a score of 53 percent [1].

One area of concern: K3’s hallucination rate climbed from 39 percent to 51 percent compared with K2.6, even as its accuracy on the AA-Omniscience Index improved from 33 percent to 46 percent [1]. The model gets more questions right, but it also fabricates more answers.

Architecture and capabilities

K3 uses a mixture-of-experts (MoE) architecture — a design that routes each input through only a subset of the model’s total parameters — activating 16 of 896 experts at a time [1]. It supports a one-million-token context window and processes images and video natively [1].

A new attention mechanism Kimi calls “Kimi Delta Attention” enables up to 6.3x faster decoding for million-token contexts, the company said, while “attention residuals” improve training efficiency by roughly 25 percent at less than 2 percent additional compute overhead [1].

Kimi positions K3’s primary use case as long-running software development with minimal human oversight, including analyzing large codebases and coordinating terminal tools across many work steps [1]. A feature the company calls “Vision in the Loop” lets the model examine screen captures, modify code, and verify visible output — targeting game development, UI design, and computer-aided design workflows [1].

Pricing and availability

Through the Kimi application programming interface (API), one million input tokens cost $0.30 with a cache hit and $3.00 without; one million output tokens cost $15.00 [1]. Those rates are a significant step up from K2.6, which is priced at $0.16 per million input tokens with a cache hit, $0.95 without, and $4.00 for output [1].

For context, Anthropic’s Sonnet 5 carries the same $3.00 input and $15.00 output pricing but scores lower on Artificial Analysis’s benchmarks, according to The Decoder’s reporting [1]. On a per-task basis, Artificial Analysis estimates K3 costs $0.94 on average, close to GPT-5.6 Sol at $1.04 and roughly half the $1.80 per task for Opus 4.8 — but far above open-weight alternatives like GLM-5.2 at $0.32 and DeepSeek V4 Pro at $0.04 [1].

K3 is already live on Kimi.com, iOS, Android, and HarmonyOS apps, the Kimi Work desktop client (version 3.1.0 and later), and Kimi Code [1]. It is also listed on OpenRouter under the identifier “moonshotai/kimi-k3,” currently served only through Moonshot itself [1]. A forthcoming platform called Kimi Hosted Agent, which will offer isolated environments for long-running tasks, is accepting waitlist sign-ups now [1].


Sources

  1. The Decoder — Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI

This article was drafted with AI from the cited sources and checked against them before publication. Spot an error? Let us know.