Benchmarks

Kimi K3 is real, GLM-5.2 is a gamble, neither is a bargain

Kimi K3 ties GPT-5.6 on the independent coding index at 4.6 times the runtime. GLM-5.2's $1.40 tokens cost $6.51 per task. The Chinese model hype, measured.

On this page
  1. Where does the hype come from?
  2. What do the independent numbers say?
  3. Is cheaper per token actually cheaper?
  4. What is Kimi K3 actually good at?
  5. Why did GLM-5.2 crash in Claude Code?
  6. How to read a coding leaderboard
  7. Where the models land in my routing

The pitch behind the new Chinese models is simple: open weights that match the Western frontier at a fraction of the price. The independent numbers are more specific. On the Coding Agent Index from Artificial Analysis, Kimi K3 does tie GPT-5.6’s medium setting, at 4.6 times the active time [1]. And GLM-5.2’s $1.40 tokens turned into $6.51 per attempted task [2].

I have not run either model against my own repositories yet, so this is not a hands-on review. It is a reading of the receipts that exist: the independent agent measurements, the arena votes, the vendors’ own tables with their own footnotes, and the license terms nobody quotes in launch threads. Short version: the capability is real, the discount mostly is not.

Where does the hype come from?

Three true facts arrived at once. Kimi K3 leads Arena’s fullstack coding leaderboard [3] and sits second on the WebDev Arena, ahead of Claude Fable 5 and GPT-5.6 [4]. Both models ship open weights at sizes nobody had opened before [5] [6]. And both list token prices far under the Western flagships [5] [7].

The arena results deserve to be taken seriously. On the fullstack board dated July 24, 2026, Kimi K3 max holds rank 1 at a rating of 1,664, ahead of GPT-5.6 Sol at 1,633 and Fable 5 at 1,623 [3]. On the WebDev Arena four days later, Claude Opus 5 max leads at 1,712 and Kimi sits second at 1,682 across 3,777 votes, above Fable 5, Sol and GLM-5.2 [4]. People with a free choice keep voting for what this model builds in a browser.

The openness is real too. Moonshot calls Kimi K3 the first open-weight model at 2.8 trillion parameters, with a 1-million-token context window, vision built in, and an API price of $3 per million input tokens and $15 for output [5]. GLM-5.2 is the smaller sibling at 753 billion parameters under a plain MIT license [6], listed at $1.40 in and $4.40 out [7]. Put those stickers next to the Western price list and the narrative writes itself.

Moonshot itself is more careful than its fans. The launch post admits K3 still trails Fable 5 and GPT-5.6 Sol, and concedes “a noticeable gap in user experience” against both [5]. That sentence, written by the vendor, turns out to be the most accurate summary of the independent data too.

What do the independent numbers say?

Artificial Analysis runs every system through the same three-part suite: 113 software engineering tasks (DeepSWE), 84 agentic terminal tasks (Terminal-Bench v2) and 124 technical codebase questions (SWE-Atlas Q&A), equally weighted into the Coding Agent Index v1.3 [8]. Pulled on July 30, 2026: Kimi K3 scores 61, exactly level with GPT-5.6 at medium effort. Opus 5 medium scores 62. GLM-5.2 scores 43 [1] [2].

One framing detail matters more than any single score. The index measures systems, not weights: Kimi ran inside its own Kimi Code CLI, GLM-5.2 ran inside Claude Code, Sol inside Codex, the Claude models inside Claude Code [1] [2]. You are never benchmarking a model. You are benchmarking a model wearing a particular agent at a particular effort setting.

Measurement Kimi K3GLM-5.2Sol mediumOpus 5 medium
Coding Agent Index 61 43 61 62 (Best value in this row)
DeepSWE, % 64 29 64 63
Terminal-Bench v2, % 84 (Best value in this row) 72 78 79
SWE-Atlas Q&A, % 37 29 40 44 (Best value in this row)
Cost per task, $ 3.18 6.51 2.99 (Best value in this row) 3.14
Active time, min 23.8 25.1 5.2 (Best value in this row) 12.2
Figure 1. Coding Agent Index v1.3, sub-scores and measured per-task costs. Artificial Analysis, pulled July 30, 2026.

The table’s story sits on the diagonal: Kimi wins the terminal, Opus wins comprehension, Sol wins the stopwatch, and GLM-5.2 wins nothing at these settings.

Look at the tie first, because it is the fairest one-line summary of Kimi. Same 61 as Sol medium, at $3.18 per task against $2.99, so about 6% more money [1]. But it needs 23.8 active minutes against Sol’s 5.2, which is the 4.6 times from the lede, and 10.6 million tokens against 5.8 million. Against Opus 5 medium the picture repeats: one index point under, four cents more per task, roughly twice the time [2].

The ceiling is still Western. Sol max reaches 67 at $7.08 per task [1]; Opus 5 xhigh matches the 67 at $8.23 and posts the best repository comprehension in the suite, 55% on SWE-Atlas Q&A [2]. Fable 5 max lands at 66 for $11.71 per task, the priciest way to buy a point that Opus xhigh also sells [2]. Kimi’s own 84% on Terminal-Bench v2 is genuine frontier territory, above Sol high’s 83% and within a few points of the max-effort flagships [1] [2].

Kimi’s soft spot is just as visible: 37% on the codebase Q&A component, under Sol medium’s 40% and far under what the Claude configurations post [1] [2]. The profile reads like an agent that is better at doing than at understanding what already exists. For greenfield work that hardly matters. For a 300,000-line brownfield repository it is the whole job.

Is cheaper per token actually cheaper?

Not by itself. A task costs tokens times token price, and an agent that needs more turns, more retries or more thinking multiplies the first factor faster than any discount shrinks the second. GLM-5.2 is the measured proof: the lowest list price on the board, yet a per-task cost above every medium-effort Western configuration [2] [7].

The list prices first, because two of them surprised me. Through August 31, 2026, Claude Sonnet 5 sells at an introductory $2 per million input tokens and $10 for output, then $3 and $15 from September 1 [9]. Kimi K3 costs $3 and $15 today [5]. The famously cheap Chinese model is 50% more expensive per token than Anthropic’s volume model this month, and list-price identical to it from autumn.

The rest of the board: Opus 5 at $5 in and $25 out, Fable 5 at $10 and $50 [9], GPT-5.6 at $5 and $30 with a long-context tier at $10 and $45 [10], and GLM-5.2 undercutting everything at $1.40 and $4.40 with $0.26 cache reads [7].

Two footnotes make even those stickers slippery. Tokenizers differ per vendor: Anthropic documents that its current tokenizer produces roughly 30% more tokens for the same text than its previous one [9], so a price per million tokens is not even a fixed unit across model families. And thinking behavior differs: Kimi K3 always reasons, the platform FAQ is explicit that it cannot be turned off, and the default effort is max [11]. You pay for tokens you never see.

list price, per million input tokens
$1.40
the cheapest model in this comparison
measured, per attempted task
$6.51
inside Claude Code, 25.1 active minutes
derived, per solved task
$15.14
cost per attempt divided by the 43% pass rate
Figure 2. What GLM-5.2's list price turns into, measured and derived. Artificial Analysis and Z.ai, July 2026.

Run that arithmetic across the board and the order changes. Sol medium comes out around $4.90 per solved task, Opus 5 medium at $5.06, Kimi K3 at $5.21, and GLM-5.2 at $15.14, three times its Western competition. The chart’s takeaway in one line: the cheapest sticker in the comparison becomes the most expensive solved change under measurement.

And that still ignores the expensive part. Ten minutes of an engineer triaging a failed run costs more than any number in that figure. The moment failures land on a human instead of a retry loop, pass rate dominates token price entirely.

What is Kimi K3 actually good at?

Terminal-heavy work, product-shaped web development, long agent runs, and every scenario where owning the weights matters. The independent suite supports it, the arenas support it, and Moonshot’s own positioning matches it [1] [3] [5]. What it is not, on current evidence, is a cheaper drop-in for repository-heavy brownfield work.

The terminal result holds up from two directions: 84% measured independently in Kimi’s own CLI [1], 88.3% in Moonshot’s preferred configuration [12]. The long-horizon claim is vendor-measured but notable: K3 tops Moonshot’s SWE-Marathon table at 42, and the same table openly footnotes that Fable 5 hit fallbacks on 35% of those tasks and that some Kimi rows ran a pre-release, H20-calibrated branch [12]. Add native vision through a dedicated encoder and the 1-million-token window [12], and Moonshot’s pitch of iterating between code and live screenshots [5] stops sounding like marketing.

The weaknesses are equally concrete. Comprehension: 37% on codebase Q&A [1]. Speed: 23.8 active minutes per task means you parallelize agents or you wait; you cannot parallelize your own attention. And “open weights” deserves napkin math before anyone plans self-hosting: 2.8 trillion parameters means roughly 1.4 TB of weights even at 4-bit quantization, before the KV cache and serving overhead. That is a cluster, not a workstation.

The paperwork matters too. The license is permissive but not boilerplate: a model-as-a-service business past $20 million in aggregate revenue needs a separate agreement with Moonshot, and any product past 100 million monthly active users or $20 million in monthly revenue must display “Kimi K3” prominently; internal use is exempt [13]. And the standard API terms let Moonshot use customer content to provide, maintain, develop, support and improve its services, with restrictions on model-training use routed to separate enterprise arrangements [14]. I would not point the public API at a proprietary repository before having that conversation in writing.

Why did GLM-5.2 crash in Claude Code?

Nobody outside Z.ai knows exactly, and that is the point. The same weights score 81% on Terminal-Bench in Z.ai’s own Terminus harness [6] and 72% inside Claude Code [2]. What collapsed in the independent run is the system around the model: tool calling, recovery, cache behavior, fit with the agent’s conventions. The measurement condemns a configuration, not necessarily the weights.

The full independent row is grim reading either way: index 43, DeepSWE 29%, codebase Q&A 29%, $6.51 per attempted task, and at 25.1 active minutes the slowest run on the whole board [2]. A trajectory that burns 25 minutes and $6.51 at a $1.40/$4.40 list price is a trajectory that spent most of its budget going in circles.

Z.ai’s own numbers come from friendlier machinery, and the model card says so plainly: SWE-Bench Pro at 62.1 through OpenHands with a tailored instruction prompt, DeepSWE with two-hour timeouts in isolated containers, Terminal-Bench through Terminus on a four-hour budget [6]. None of that is cheating. It is a vendor showing the model in the environment it was tuned for. It just is not evidence about the agent you actually run.

There is still a rational lane for it. At 753 billion parameters under MIT [6], GLM-5.2 is the realistic self-hosting candidate of the two, and at $1.40 per million tokens it is a sensible pilot for bulk work with automatic verification: codemods with test suites, migration candidates, anything where a deterministic check catches the 57% that fails and failures cost only compute. The moment a human reviews the failures, the arithmetic flips back.

How to read a coding leaderboard

The Kimi and GLM launches are a case study in benchmark literacy, so here is the map I use before believing any number.

SignalWhat it rewardsWhat it cannot tell you
DeepSWE, 113 taskscompleting end-to-end engineering tasksfit with your codebase’s conventions
Terminal-Bench v2, 84 tasksdriving a shell to a verified end statecomprehension of a large existing repository
SWE-Atlas Q&A, 124 tasksanswering technical questions about codethe ability to land the change it describes
WebDev and Fullstack Arenawhat people prefer in head-to-head votescorrectness, tests, security, maintainability
Vendor model cardsthe model at its best, in its own setupcomparability between rows

Then apply three discounts. First, infrastructure noise: Anthropic’s engineering team measured a 6-percentage-point swing on Terminal-Bench 2.0 from resource limits alone, watched infrastructure error rates fall from 5.8% to 0.5% as limits loosened, and concluded that leaderboard gaps under about 3 points deserve skepticism until the configurations are documented and matched [15]. Kimi’s tie with Sol medium sits inside that band; treat it as parity, not as either model winning.

Second, footnote asymmetry. Moonshot’s comparison table runs each competitor in a different harness, includes fallback models on some rows, and scores its own in-house benchmark, where Fable 5 logged 13 fallbacks and one refusal across 80 tasks [12]. Credit where due: Moonshot printed those footnotes itself. The launch threads quoting the table did not.

Third, version discipline. The independent index is versioned (v1.3 today) because its task mix changes [8]; a score from one version is not comparable to a score from another, no matter how similar the name looks. Any comparison that does not state harness, effort, budget and version is a vibe wearing a decimal point.

Where the models land in my routing

Everything above changes what I would trial first, not what I would standardize on. My routing table today, with the two newcomers slotted honestly:

The workMy pick today
Routine changes with solid testsSonnet 5, at $2/$10 until August 31, 2026
Fast agent loops and terminal workGPT-5.6 at medium or high effort
Brownfield work in a large repositoryOpus 5 at medium or high effort
Greenfield fullstack and UI-heavy buildsKimi K3, the lane the arena votes back
Bulk work with automatic verificationGLM-5.2 in its native harness, as a pilot
Escalation when everything else failsFable 5, sparingly

It is the same arithmetic that Opus 5’s launch week taught me: what you actually buy is approved changes, and everything else, tokens included, is an input.

Two results would move these rows. A Kimi K3 run inside a neutral harness (Claude Code or Codex) holding that 61, which would prove the score belongs to the model rather than to its home-field CLI. And an independent GLM-5.2 measurement in its native agent, run by someone who does not sell it. Until then the honest summary stands: the models are real, the prices are marketing, and the receipts are above.

Sources

  1. Codex vs Kimi Code CLI: coding agent comparisonArtificial Analysis
  2. Claude Code vs OpenCode: coding agent comparisonArtificial Analysis
  3. Fullstack Arena leaderboardArena · 2026-07-24
  4. WebDev Arena leaderboardArena · 2026-07-28
  5. Kimi K3Moonshot AI
  6. GLM-5.2 model cardZ.ai on Hugging Face
  7. GLM-5.2 API pricingZ.ai Docs
  8. Coding Agent IndexArtificial Analysis
  9. PricingClaude Platform Docs
  10. API pricingOpenAI Developers
  11. Kimi K3 quickstartKimi Platform Docs
  12. Kimi K3 model cardMoonshot AI on Hugging Face
  13. Kimi K3 licenseMoonshot AI on GitHub
  14. Model use agreementKimi Platform Docs
  15. Benchmark scores and infrastructure noiseAnthropic Engineering