Benchmark

Kimi K3: Cost and Performance Edge

Kimi K3 offers strong performance at a fraction of the cost of leading American models, excelling in frontend work while earning praise from users.

September 24, 20251 min readKimi K3costperformanceAI models

Models like Kimi K3 from Moonshot AI in China have put into perspective how costly American models are for a relatively equally intelligent piece of work. At artificialanalysis.ai, Kimi K3 scores 57, placing third behind Fable 5 and GPT-5.6.

57 Kimi K3 score

And its ~two thirds cheaper than Fable 5 with respect to cost per task.

~66% cheaper than Fable 5 cost per task

I am personally aiming to shift my monthly subscription to Kimi K3 next month when my Claude subscription ends. According to X sentiment, those who are using Kimi consistently have nothing but praises for it. Hopefully Moonshot AI will have built up the capacity needed by the time I'm getting to it.

Btw Kimi is also best in class in frontend work by a big margin.

best in class Frontend work Kimi K3

Another model recently released from China is Qwen 3.8. With more than 2T parameters this model is projected to match Kimi or even surpass it. But this model has not been officially released yet, just some previews.

American models though are not to be left behind. Grok 4.5 is actually the best model in terms of intelligence, speed and cost per task. If you were to draw a map that takes into factor the three above, I bet Grok 4.5 would be the best.

Gemini 3.6 Flash also recently released gives the best speed you'd require in an interactive app. We use it in our chat interfaces exclusively.

50 Gemini 3.6 Flash score Second to GPT-OSS-120b (24)

As you can see above, Gemini 3.6 Flash is second to GPT-OSS-120b which, despite being fast, is incredibly dumb with a score of 24, compared to Flash's 50.