Direct Connection Test
Connects directly from your browser to test live streaming speed and latency without proxy delay.
Serialized Telemetry Agent Protocol (1-Paste)
Copy the serialized benchmark script once. Paste into your AI (ChatGPT, Grok, Claude, or Terminal). The AI executes the sub-command, benchmarks its own speed, and submits the results directly back to this page.
Copy Serialized Benchmark Script
Generates a cryptographically tied session script. Paste it into your AI chat window and hit Enter. You don't need to copy anything back—the AI transmits its telemetry back to this serial number.
"This AI's harness has been constructed so as to hide our ability to test it easily. Instead, you can paste the following into a fresh session's prompt window, and use a timer to check the speed."
📊 Speedtest Timing Worksheet
Human-Audited VerificationAI Speed Diagnostics: Why Your Models Are Slowing Down
Independent empirical telemetry, GPU cluster bottleneck analysis, and silent model downgrade detection across Claude Fable 5, Gemini 3.7 Flash, Grok 4.6, GPT-5.6, and DeepSeek V4.
Why is Claude Fable 5 & Opus 5 responding so slowly?
Claude Fable 5 and Opus 5 represent the apex of deep agentic reasoning and nuanced code synthesis. However, under high cluster demand, Anthropic's autoregressive reasoning pipeline suffers severe KV-cache expansion overhead. During peak hours, standard interactive token streams frequently throttle down to 6 to 15 tok/s. To bypass this, Anthropic offers an accelerated "Fast Mode" running on dedicated hardware at twice the pricing tier.
Are frontier AI labs choking token flows or coprocessing with inferior models?
Yes. The AI industry is in an unprecedented compute squeeze. Running hundreds of billions of parameters on full precision is cost-prohibitive. When GPU cluster capacity is saturated, providers deploy dynamic throttling (choking token delivery) or silently route queries to smaller distilled draft models or older model checkpoints. Slowtest's Line 1 Identity Verification mandates that the model declare its exact architecture before answering, exposing covert fallback routing.
How fast is Google Gemini 3.7 Flash compared to other frontier models?
Google's Gemini 3.7 Flash runs natively on custom Google TPU v6 infrastructure with optimized speculative decoding. While deliberate reasoning giants like Claude Opus 5 average 6–18 tok/s, Gemini 3.7 Flash delivers sustained throughput of 160 to 210+ tok/s with sub-400ms time-to-first-token (TTFT) and an expansive 2M+ token context window.
What is tokens per second (tok/s) and what is a normal score in 2026?
Tokens per second (tok/s) measures the raw text generation throughput of an AI model:
• < 15 tok/s (Snail / Throttled): Indicates severe cluster load, token choking, or deep multi-step ponder overhead.
• 15 – 45 tok/s (Standard): Conversational interactive baseline.
• 45 – 90 tok/s (Fast Agentic): High-performance production speed (Grok 4.6, GPT-5.6 Sol standard).
• 100 – 250+ tok/s (Blazing Edge): TPU/LPU-accelerated hosted infrastructure (Gemini 3.7 Flash, Cerebras, Groq).
Why should I test open models on scorching fast hosted hardware?
Proprietary frontier models often lock users behind heavy queue firewalls and rate limits. Modern open-weights models like DeepSeek V4, Qwen 3.8, and GLM-4 running on edge-routed inference platforms (like Abacus.AI, Monica, and Sider) deliver frontier-level intelligence at 3x to 10x the velocity without arbitrary token caps.