Leaderboard: coding index and price
Eleven models were evaluated using LiveCodeBench, which employs new tasks to minimise training-data contamination. The coding index reflects success on isolated generation, refactoring and debugging tasks. GPT-5.6 Sol leads with an index of 78.3 at 11.25 USD per million tokens. Claude Opus 5 follows at 78 for 10 USD. GPT-5.6 Terra scores 76.7 for 4.5 USD. Claude Fable 5 reaches 76.5 but costs 20 USD, the highest in the set. Kimi K3 delivers 76.2 for 6 USD. Grok 4.5 scores 72.4 at 3 USD. Muse Spark 1.1 achieves 71.3 for 2 USD. Gemini 3.5 Flash records 70.1 at 3.38 USD. Gemini 3.6 Flash scores 69.2 for 3 USD. DeepSeek V4 Flash 0731 posts 69.1 at 0.17 USD. Qwen3.7 Max trails at 66 for 3.75 USD. All figures reflect the August 2026 snapshot.
What the benchmarks miss
The coding index measures performance on isolated tasks, not the ability to maintain context across a large repository, resolve dependencies or debug iteratively. LiveCodeBench reduces contamination through novel tasks, yet production code demands architectural understanding, test suites and CI/CD integration. Token price excludes latency, context-window limits and tooling overhead. A model scoring 78.3 may not prove more effective in production than one scoring 71.3 if the latter makes better use of tools and shorter contexts. The gap between benchmark scores and real-world utility widens when workflows require multi-file edits, dependency resolution or continuous integration feedback loops that no single-turn benchmark captures.
How to choose for production use
Organisations with tight budgets will find DeepSeek V4 Flash 0731 the most cost-effective option at 0.17 USD per million tokens with an index of 69.1. Where maximum generation quality is the priority, GPT-5.6 Sol at 11.25 USD and 78.3 index is justified. The sweet spot lies with GPT-5.6 Terra at 4.5 USD and 76.7, or Muse Spark 1.1 at 2 USD and 71.3. Avoid defaulting to the most expensive model: Claude Fable 5 at 20 USD does not deliver proportional gains over GPT-5.6 Terra. The decision should weigh token cost against the engineering time saved by higher first-pass accuracy, the size of the codebase the model must navigate, and the tooling ecosystem each provider supports.
Frequently asked questions
Which AI model codes best according to LiveCodeBench?
The highest coding index of 78.3 belongs to GPT-5.6 Sol, closely followed by Claude Opus 5 at 78. Both are from OpenAI and Anthropic. Third place GPT-5.6 Terra achieves 76.7. Differences at the top are minimal, but prices differ significantly.
Is it worth paying more for a pricier model when programming?
Not always. Claude Fable 5 costs 20 USD and achieves 76.5, while GPT-5.6 Terra at 4.5 USD achieves 76.7. DeepSeek V4 Flash 0731 offers 69.1 for 0.17 USD. A pricier model makes sense only for critical tasks requiring maximum generation accuracy.
How reliable are coding benchmarks for model selection?
LiveCodeBench uses new tasks to prevent contamination, but measures isolated tasks. Real code work requires repository context, debugging, tests, and tool integration. An index of 78.3 versus 71.3 does not mean a direct transfer to productivity. Test models on your own tasks before deployment.
DATA SOURCES AND METHOD
Všechna čísla v tomto srovnání pocházejí z uvedených zdrojů k datu 4 August 2026. Grafy generuje institut CIAD přímo ze zdrojových dat; textová analýza čísla nikdy nedopočítává ani neodhaduje.