Kimi K3 and Opus 5 on Cogny Bench: A Bigger Open Model Isn't a Better Buy
Kimi K3 and Opus 5 on Cogny Bench: A Bigger Open Model Isn't a Better Buy
Berget — the European, open-weights inference platform we run our sovereign-friendly options on — added Kimi K3 (Moonshot's 2.8-trillion-parameter flagship) to its catalog on 2026-07-28. So we did the obvious thing and pointed Cogny Bench at it, right next to the model it's meant to succeed on Berget, GLM-5.2.
The same sweep picked up Claude Opus 5, which walked straight to the top of the board. Two findings, one run.
The numbers
A representative slice of the board (the chart has the full field):
| Model | Avg score | $/run |
|---|---|---|
| Opus 5 | 94.0 | $0.320 |
| Gemini 3.5 Flash | 93.6 | $0.153 |
| Sonnet 5 — intro ($2/$10) | 93.6 | $0.098 |
| Opus 4.7 | 93.2 | $0.489 |
| Opus 4.8 | 93.1 | $0.657 |
| GPT-5.6 Terra | 93.1 | $0.052 |
| Kimi K3 (Berget) | 90.2 | $0.201 |
| GLM-5.2 (Berget) | 90.0 | $0.075 |
| Fable 5 | 86.0 | $0.390 |
| Haiku 4.5 | 73.8 | $0.108 |
Scores are the combined metric — deterministic answer checks blended with a fixed LLM judge — averaged across nine marketing-analytics reasoning problems spanning confounds, mix shifts, tracking outages, currency conversion and forecasting. Single runs (n=1) per cell, so expect a few points of wobble. One grader cell currently mis-scores for every model and drags all averages down by the same few points — it doesn't change the ranking, but it's why the top scores sit in the mid-90s rather than higher.
Pricing is real, from each provider's authoritative source: Kimi K3 and GLM-5.2 from Berget's published rates (€3 / €15 and €1.4 / €4.4 per million tokens respectively), Opus 5 from Anthropic's list ($5 / $25).
Kimi K3: frontier reasoning, one expensive stumble
On eight of the nine problems Kimi K3 is genuinely frontier — 98s and 99s across the mix-shift, tracking-outage, currency and forecasting traps. Its analysis is not the problem.
Where it breaks is the same place the base-tier models break: a single confounded-comparison problem, where two effects overlap and you have to isolate the real one against a control the data hides in plain sight. Kimi K3 over-attributes — and it does so expensively, thrashing through dozens of tool calls and roughly $1.10 on that one problem before committing to the wrong read. That single cell, plus the shared grader artifact, is the whole gap between its 90.2 and the near-99s it posts everywhere else.
The value question: Kimi K3 vs GLM-5.2
Here's the part that actually changes a buying decision. On Berget:
| Kimi K3 | GLM-5.2 | |
|---|---|---|
| Avg score | 90.2 | 90.0 |
| $/run (bench) | $0.201 | $0.075 |
| Berget list (per M tokens) | €3 / €15 | €1.4 / €4.4 |
| Parameters | 2.8T | 753B |
Kimi K3 is the bigger, newer, more expensive model — and on our marketing-analytics reasoning it scores within two tenths of a point of GLM-5.2 while costing ~2.7× as much per run. Both handle the routine problems cleanly; both trip the same confound. On this workload, "bigger open model" and "better buy" are not the same question — GLM-5.2 remains the value pick on Berget, and Kimi K3's extra capacity doesn't show up where our reports need it.
That's exactly why Cogny doesn't hardcode one model. Our managed Cogny mix picks a model per job, and the bench is one of the inputs to that choice. For the open-weights, EU-sovereign lane, the bench says default to GLM-5.2 and keep Kimi K3 on the shelf for the workloads (long-context, vision) where its extra size might actually earn its price.
Opus 5 takes the top of the board — and made our grader more honest
In the same sweep, Claude Opus 5 posted 94.0, the first model to clear 94 on our board. It read cleanly through the two problems that discriminate models best — the confounded comparison Kimi tripped on, and the forecasting trap — and only the shared grader artifact keeps it out of the high 90s.
Getting that number honest took a fix worth reporting rather than hiding. On one problem, Opus 5's answer was completely correct and the judge scored its reasoning near-perfect — but our deterministic grader initially read it as a zero, because Opus 5 returned its structured answer in a serialization our grader didn't expect and the field checks found nothing. That's a well-known tool-call formatting quirk, not a wrong answer. We taught the grader to accept the alternative serialization, re-scored the cell, and it landed the near-perfect mark it always deserved. The grader's sanity checks still pass cleanly, so the fix made the harness more robust without loosening it.
The takeaway
Two lessons from one afternoon's run, both the kind the bench exists to catch:
- For open weights on Berget, size isn't the buy signal. Kimi K3 costs 2.7× GLM-5.2 for a statistically identical score on our workload. Match the model to the job — kept honest by the eval — instead of reaching for the biggest one.
- For the frontier, Opus 5 is the new ceiling on Cogny Bench at 94.0, and the exercise made our grader more robust in the process.
Want the methodology behind the traps? Read how Cogny Bench is built — the categories of reasoning we test, and why we keep the problems themselves private.