What the data shows about model speed

Gemini 3.5 Flash leads token throughput at 267.8 tokens per second, followed by Gemini 3.6 Flash at 230.4 and Qwen3.7 Max at 204.1. At the bottom sit Kimi K3 with 34.5 tokens per second, Grok 4.5 with 55.6 and Claude Opus 5 with 59.1. For time to first token (TTFT), GLM-5.2 is fastest at 0.96 seconds, ahead of MiniMax-M3 at 1.38 seconds and Qwen3.7 Max at 1.49 seconds. The slowest are GPT-5.6 Terra at 82.52 seconds and GPT-5.6 Sol at 60.18 seconds. The intelligence index is topped by Claude Opus 5 at 60.7, then Claude Fable 5 at 59.9 and GPT-5.6 Sol at 58.9. Muse Spark 1.1 stands out with an intelligence index of 50.6 while delivering 172.2 tokens per second and a TTFT of 2.46 seconds.

Critical interpretation of the trade-offs

The numbers expose a clear trade-off. Models with the highest intelligence index, Claude Opus 5, Claude Fable 5 and GPT-5.6 Sol, provide throughput an order of magnitude lower and latency an order of magnitude higher than the fast Gemini Flash or GLM-5.2 models. A TTFT spread of 0.96 to 82.52 seconds means users of the slowest models wait over a minute for the first token, destroying the interactive chat experience. For batch processing, throughput is decisive: Gemini 3.5 Flash outperforms the slowest model by a factor of 7.8. Laboratory benchmarks do not capture variability under real load or infrastructure cost.

Practical implications for deployment

Organisations deploying chatbots should prioritise TTFT: GLM-5.2, MiniMax-M3 and Qwen3.7 Max deliver sub-two-second responses at acceptable intelligence levels. Agents that require long contexts and rapid generation benefit from the throughput of Gemini 3.5 Flash or Muse Spark 1.1. Offline batch jobs favour maximum throughput regardless of latency. Model selection must acknowledge that high intelligence, Claude Opus 5 at 60.7, comes at a cost of 34.11 seconds latency and 59.1 tokens per second throughput.

Frequently asked questions

Which model has the lowest TTFT latency?

GLM-5.2 achieves the lowest TTFT latency at 0.96 s, followed by MiniMax-M3 at 1.38 s and Qwen3.7 Max at 1.49 s according to measurements from August 2026.

How large is the throughput difference between the fastest and slowest model?

The token throughput difference between the fastest model Gemini 3.5 Flash and the slowest Kimi K3 is 7.8× according to the precomputed ratio in the data.

Which model offers the best balance between speed and intelligence?

Muse Spark 1.1 achieves an intelligence index of 50.6 at a throughput of 172.2 tokens/s and TTFT of 2.46 s, which according to the table represents a strong balance of both metrics.

DATA SOURCES AND METHOD

Všechna čísla v tomto srovnání pocházejí z uvedených zdrojů k datu 3 August 2026. Grafy generuje institut CIAD přímo ze zdrojových dat; textová analýza čísla nikdy nedopočítává ani neodhaduje.