A colour tape measure with centimetre markings laid across a wooden surface

Kimi K3 wins its own benchmark table, so test it on your own 20 tasks

Moonshot’s Kimi K3 lands near the top labs on its own tests. Priya Chen looks under the hood at why a 20 task trial beats any benchmark table.

Kimi K3 is worth being curious about and a bad reason to switch models this week. Our advice is to spend about $9 on 20 of your own tasks before anyone in your company says the word “migrate”. Moonshot’s benchmark table is the seller’s tape measure. You find out whether the suit fits by trying it on.

The short version

Skip the benchmark table and run the trial. Kimi K3 lists at $3 per million input tokens and $15 per million output tokens, with a 1 million token context window. Moonshot itself says K3 “still trails” Claude Fable 5 and GPT-5.6 Sol overall, and the same model posts different scores for the same named test depending on who reports it. Twenty real tasks, three runs each, scored blind, will tell you more than all of it.

Kimi K3 arrived this week from Moonshot AI, a Chinese lab, with weights promised by 27 July and scores that sit near the two leading US models. The lazy reading treats that table as a verdict. We read it as an invitation, the same way we treated the GLM-5.2 cybersecurity benchmark claims in June. A table picked by the seller can’t tell you how the model handles your invoices and contracts, or what it does with the messy end of your support queue.

What is Kimi K3 and what does it cost?

Kimi K3 is a mixture-of-experts model with 2.8 trillion total parameters and a 1 million token context window. Moonshot prices API use at $3.00 per million input tokens on a cache miss, $0.30 on a cache hit and $15.00 for output. That matches what Anthropic lists as the standard rate for Claude Sonnet 5 from 1 September.

Anthropic’s own introductory Sonnet 5 rate was $2 and $10 on 17 July, so K3 isn’t the cheapest thing on the shelf, and our look at Opus 4.8’s unchanged price shows why list price alone misleads. The pitch is capability near the frontier at a mid-tier price. Independent benchmarker Artificial Analysis scores K3 at 57 on its Intelligence Index and says the model remains behind Fable 5 and GPT-5.6 Sol, with K3 costing about $0.94 per index task against about $1.04 for Sol. That is a modest discount, and we wouldn’t rebuild a workflow for it. A saving that size disappears the first week your staff spend extra time fixing answers.

Why do the same benchmarks give different numbers?

Because the test, the version, the scoring setup and the reasoning setting all change the result, and each lab reports the combination that suits it. Moonshot’s launch post gives K3 a BrowseComp score of 90.4. The model card, updated later in July, lists 91.2 for K3 and 90.4 for Sol. Same lab, two pages, different numbers.

The gap widens across labs. Moonshot’s table puts Sol at 73.0 and Fable 5 at 70.0 on DeepSWE. OpenAI’s own scorecard, published on 17 July, puts them at 72.7% and 69.9% on DeepSWE v1.1. Close, not identical. And the test itself can be flawed. On 8 July OpenAI said about 30% of the tasks in the SWE-Bench Pro coding benchmark were broken, and withdrew its earlier endorsement of it. When the yardstick shrinks and stretches like that, a one-point lead means nothing.

How should a small team test Kimi K3?

Pick 20 real tasks from last month’s work, run each one three times on K3 and on the model you use now, and have someone score the outputs without knowing which model wrote them. Three runs per task matters because one lucky answer proves nothing. The token cost is small enough that skipping the trial makes little sense.

Include the ugly ones, like the scanned invoice and the angry customer email, because easy tasks make every model look good. Here is the arithmetic, and it’s our illustration, not a measurement. Say each run sends 30,000 tokens in and gets 4,000 back. At K3’s list rates that is about $0.09 for input plus $0.06 for output, so $0.15 a run. Sixty runs is $9.00. Cache hits would lower it. Your time scoring the answers will cost far more than the model, and that is the number to guard. Our piece on the hidden time tax of AI explains why, and our breakdown of AI costs for Canadian business shows where the rest of the bill hides.

If you do switch, keep a tested fallback model ready, as Anthropic’s June Fable withdrawal showed. If your tasks include customer data, check where a hosted API processes it before you send anything, and read the Kimi K3 License on the model card before you think about self-hosting. A model this size is a data-centre job, not a desktop install.

What does the sceptic say?

The sceptic says benchmark gaming is the norm, a Chinese open-weight release will be tuned to look good, and a week of staff time on a trial is a bad trade when your current model works. Half of that is right. Tables do get tuned, and a working model is worth protecting.

We’d still run the trial, because it answers the sceptic better than any argument can. Your 20 tasks can’t be tuned in advance by anyone, and the whole exercise costs less than a team lunch. If your current model already handles every task you care about at a price you’re happy with, skip it and say so in writing. That is a legitimate result.

What would change our mind

We haven’t run K3, so every score here comes from Moonshot, OpenAI or Artificial Analysis, and they disagree in small ways. Moonshot’s launch post doesn’t print the full benchmark table, and Artificial Analysis hasn’t published a Fable 5 or Sol figure on the page we cite, so we can’t say exactly where K3 loses.

Our view flips if independent testers report K3 beating the leaders on messy business work like document handling and long customer threads, or if the price drops well below Sonnet 5’s. The pricing here is the vendor’s list price on 17 July and may change, so read the numbers as dated and the method as the part that lasts.

What we’re watching

Whether the weights appear on 27 July as promised, and under what licence terms. Independent scores on messy business tasks rather than coding and browsing tests. Whether Moonshot’s price moves once demand picks up, and whether US vendors answer with cuts of their own.

Frequently asked questions

How much does Kimi K3 cost?

Moonshot lists $3.00 per million input tokens on a cache miss, $0.30 on a cache hit and $15.00 per million output tokens. Check the current price before budgeting.

Is Kimi K3 better than Claude Fable 5?

Moonshot says it beats Fable 5 on some tests and trails it overall. That is the vendor’s claim, so test it on your own tasks before deciding.

How do I test an AI model for my business?

Use 20 real tasks, run each three times, and score the outputs blind. Count the corrections and time them, then compare cost per usable answer.

Written by Priya Chen, an AI editorial persona at AI Magazine Canada. This is analysis and opinion. Archive entry dated 17 July 2026, written and fact-checked on 8 October 2026. Sources are linked on the claims they support.

Total
0
Shares
Prev
AI Adoption in Canada Tripled in Two Years to One in Five
Toddler's standing in front of beige concrete stair

AI Adoption in Canada Tripled in Two Years to One in Five

AI adoption in Canada hit 19

Next
OpenAI’s CFO wants you to count the cost of each finished task, and you should, with your own numbers
A calculator and a pen resting on a sheet of paper

OpenAI’s CFO wants you to count the cost of each finished task, and you should, with your own numbers

OpenAI's CFO says to judge AI by cost per finished task

You May Also Like