Kimi K3 Sets New Benchmark in Large Language Models

Moonshot AI has unveiled Kimi K3, a 2.8 trillion parameter model that’s quickly gaining recognition across various benchmarks. The open-weight release is slated for July 27, 2026.

The new model demonstrates impressive capabilities, rivaling established leaders like Claude Opus and GPT models while offering competitive pricing at $3 per million input tokens and $15 per million output tokens—the highest among Chinese AI labs to date.

Performance Highlights

  • Achieved an Elo of 1547 on long-horizon knowledge work evaluations, surpassing Kimi K2.6 by +732 points
  • Cost per task is $0.94, similar to GPT-5.6 Sol ($1.04) and lower than Opus 4.8 ($1.80)
  • Significantly reduced token usage compared to previous versions (21% fewer output tokens than K2.6)
  • Currently leading on Arena.ai’s Frontend Code arena, surpassing even Claude Fable 5

Visual Capabilities Test

I tested Kimi K3 by requesting an SVG of a pelican riding a bicycle—a benchmark I initially created as a lighthearted challenge for LLMs.

The model generated the image in just 95 input tokens and 16,658 output tokens (with 13,241 used for reasoning) at a cost of only 25 cents!

When prompted to describe its own creation using my alt-text prompt, Kimi K3 accurately identified key elements including: the pelican’s white feathers and orange beak, the red scarf and bicycle, motion lines indicating movement, and even subtle details like the yellow sun and tiny flowers in the background—all for just 0.6 cents.

This performance suggests that while my “pelican test” may have outlived its usefulness as a serious benchmark, it still offers valuable insights into how these models process visual information and generate creative content.