Overview of the September 2026 benchmark
Companies deploying language models in September 2026 face a choice between response speed and output quality. A new comparison covers twelve models from Google, OpenAI, Meta, Z AI, Anthropic, xAI, Alibaba and Kimi across three metrics: median tokens per second, median time to first token (TTFT) and an intelligence index. The data reveal a wide spread, from Gemini 3.7 Flash at 307.1 tokens per second down to models with substantially lower throughput, and a TTFT range from 1.41 to 244.37 seconds.
Speed leaders and laggards
By median tokens per second, Gemini 3.7 Flash leads at 307.1, followed by GPT-5.6 Terra at 105.7 and Muse Spark 1.2 at 77.3. Next come GPT-5.6 Sol at 76.7, GLM-5.3 at 73.4 and Claude Fable 5.1 at 67.6 tokens per second. At the other end sit Qwen3.8 2.4T A95B, Qwen3.8 Max and Kimi K3, each at 39.5 tokens per second.
The TTFT spread runs from 1.41 to 244.37 seconds. Qwen3.8 Max posts the lowest at 1.41 seconds, while Claude Fable 5.1 records the highest at 244.37 seconds. That same model carries an intelligence index of 65.7 alongside its markedly higher first-token latency. GLM-5.3, GLM-5.3-Flash and Qwen3.8 Max all keep TTFT under two seconds at 1.45, 1.43 and 1.41 seconds respectively.
Gemini 3.7 Flash, the throughput leader, sits in the middle on TTFT at 7 seconds. GPT-5.6 Terra and GPT-5.6 Sol show TTFTs of 94.25 and 101.03 seconds, illustrating that high or medium token throughput and low first-token latency do not move together in this data set.
What speed does not tell about quality
TTFT measures how long a user waits for the first character of a response, and for conversational deployments such as a customer-support chatbot it is often more critical than total token throughput. Qwen3.8 Max with a TTFT of 1.41 seconds can begin answering almost immediately, whereas Claude Fable 5.1 at 244.37 seconds represents a wait that would be obvious in an interactive interface. For batch processing of large text volumes, where the user waits for the whole job rather than the first character, median tokens per second matters more, and Gemini 3.7 Flash leads at 307.1.
Intelligence index values vary independently: Claude Opus 5 scores 63.1 at 48.4 tokens per second, GPT-5.6 Sol scores 60.9 at 76.7 tokens per second, and Gemini 3.7 Flash scores 56 at 307.1 tokens per second. The table gives no reason to link high speed with a high intelligence index or vice versa.
All figures are medians measured under defined conditions. Real-world operation with custom prompts, context lengths or API infrastructure load may differ. A median describes a typical value, not the behaviour of every single query, which must be kept in mind when interpreting the results.
Choosing a model by deployment type
For firms building a chatbot or AI agent with interactive response, TTFT is the relevant metric from this table. Models such as GLM-5.3 at 1.45 seconds, GLM-5.3-Flash at 1.43 seconds or Qwen3.8 Max at 1.41 seconds offer a practically immediate first reply, which feels like a smooth conversation even with lower total token throughput.
For batch processing, summarising large documents or bulk content generation, median tokens per second is the priority, where Gemini 3.7 Flash leads at 307.1. Selecting solely by intelligence index, such as the 65.7 of Claude Fable 5.1, can mean accepting a TTFT of 244.37 seconds, a trade-off that may not suit interactive deployment.
Frequently asked questions
Which AI model is the fastest in the table according to tokens per second?
According to the median tokens per second, Gemini 3.7 Flash leads in the table with a value of 307.1 tokens per second from Google. Second in order is GPT-5.6 Terra from OpenAI with 105.7 tokens per second. The ratio of the fastest to the slowest model in the table is approximately 7.8×.
What does TTFT mean and why is this metric important?
TTFT denotes the median time to the first response token. In the table this value ranges from 1.41 to 244.37 seconds. The lowest TTFT of 1.41 seconds belongs to Qwen3.8 Max, the highest 244.37 seconds to Claude Fable 5.1. For chatbots and interactive applications this is a key indicator of conversation fluidity.
Does model speed correlate with its intelligence index?
According to the given data, not unequivocally. Gemini 3.7 Flash has the highest speed of 307.1 tokens per second and an intelligence index of 56, while Claude Fable 5.1 with a lower speed of 67.6 tokens per second achieves an intelligence index of 65.7. Speed and intelligence index evolve independently of each other in the table.
DATA SOURCES AND METHOD
Všechna čísla v tomto srovnání pocházejí z uvedených zdrojů k datu 2 September 2026. Grafy generuje institut CIAD přímo ze zdrojových dat; textová analýza čísla nikdy nedopočítává ani neodhaduje.