Summary Performance and Pricing

Claude Opus 5 from Anthropic and Kimi K3 from Kimi were compared using official benchmark data from July 2026. The evaluation covers aggregate intelligence and coding indices, deep benchmarks for science and long context, as well as price and throughput.

Claude Opus 5 scores 60.7 on the aggregate intelligence index and 78 on the coding index. Kimi K3 records 57.1 on intelligence and 76.2 on coding. The intelligence gap is 3.6 points in favour of Claude Opus 5.

Pricing per million tokens is 10 USD for Claude Opus 5 and 6 USD for Kimi K3. Throughput measures 61.4 tokens per second for Claude Opus 5 and 32.9 tokens per second for Kimi K3. Time to first token (TTFT) is 32.53 seconds for Claude Opus 5 and 5.48 seconds for Kimi K3.

Deep Benchmark Results

On the Humanity’s Last Exam (HLE) benchmark, which tests complex cross-disciplinary questions, Claude Opus 5 achieves 52.6 percent versus 44.3 percent for Kimi K3, a lead of 8.3 percentage points. This benchmark measures the ability to answer expert-level questions across scientific domains and is a key indicator for research workloads.

On GPQA Diamond, a graduate-level science question set, Kimi K3 edges ahead with 93.5 percent against 93.2 percent for Claude Opus 5. On SciCode, which evaluates scientific algorithm implementation, Kimi K3 scores 58.7 percent to 55.7 percent. On Long Context Reasoning (LCR), Kimi K3 reaches 74.7 percent compared with 70 percent for Claude Opus 5.

Implications for Enterprise Deployment

For workloads that demand deep reasoning and complex judgement, particularly in scientific research and multi-step expert analysis, Claude Opus 5 holds an advantage. Its lead on HLE indicates stronger performance on multi-disciplinary reasoning that mirrors real-world research tasks.

For applications involving long documents, scientific code generation, or cost-sensitive autonomous agents, Kimi K3 delivers comparable or superior results at lower cost and with faster first-token latency. The lower per-token price and quicker TTFT make it suitable for high-volume or interactive terminal-based workflows where responsiveness matters.

The choice depends on whether the priority is peak reasoning quality or throughput and cost efficiency at near-parity on coding and long-context tasks. Organisations running large-scale agent pipelines will find Kimi K3’s economics compelling, while research teams needing the highest reasoning fidelity may prefer Claude Opus 5 despite the premium.

Frequently asked questions

Which model is better for scientific tasks?

Claude Opus 5 leads in expert questions across disciplines (HLE) at 52.6%, while Kimi K3 achieves 44.3%. In PhD-level scientific questions (GPQA Diamond), however, Kimi K3 narrowly leads Claude Opus 5 with 93.5% versus 93.2%.

What is the price difference between the two models?

Claude Opus 5 costs 10 USD per million tokens, while Kimi K3 costs 6 USD. The calculated price ratio is 1.7x in favor of Kimi K3.

Which model is faster?

Kimi K3 achieves a speed of 32.9 tokens per second, while Claude Opus 5 achieves 61.4 tokens per second. In time-to-first-token (TTFT), however, Kimi K3 is significantly faster at 5.48 seconds versus 32.53 seconds for Claude Opus 5.

DATA SOURCES AND METHOD

Všechna čísla v tomto srovnání pocházejí z uvedených zdrojů k datu 29 July 2026. Grafy generuje institut CIAD přímo ze zdrojových dat; textová analýza čísla nikdy nedopočítává ani neodhaduje.