Choose your Test Mode:
0 50 100 200 400+
0
TOKENS / SEC
READY FOR TEST
EDGE PING
-- ms
Internet latency to edge
VELOCITY
-- tok/s
Generation throughput
INFERENCE TIME
-- s
Pure AI compute latency
FIDELITY
-- %
Format & negative constraints

Direct Connection Test

Connects directly from your browser to test live streaming speed and latency without proxy delay.

Serialized Telemetry Agent Protocol (1-Paste)

Copy the serialized benchmark script once. Paste into your AI (ChatGPT, Grok, Claude, or Terminal). The AI executes the sub-command, benchmarks its own speed, and submits the results directly back to this page.

Edge Ping: Measuring...
Active Serial: SLOW-READY
Listener: READY
THE 1-PASTE PROTOCOL

Copy Serialized Benchmark Script

Generates a cryptographically tied session script. Paste it into your AI chat window and hit Enter. You don't need to copy anything back—the AI transmits its telemetry back to this serial number.

Ready to Benchmark
Click below to generate and copy your serialized benchmark script
⚠️ HARNESS RESTRICTION DETECTED Sandbox Egress Firewall Fallback

"This AI's harness has been constructed so as to hide our ability to test it easily. Instead, you can paste the following into a fresh session's prompt window, and use a timer to check the speed."

Standardized Fresh Session Benchmark Prompt:
[SLOWTEST AUDIT & IDENTITY VERIFICATION] 1. LINE 1: Declare your exact model name, version, and architecture (format: MODEL: <name>). 2. REASONING: Deduce the next number in this sequence and explain why in 1 sentence: 2, 6, 12, 20, 30, ? 3. SYNTHESIS: Explain why subsea fiber-optic cables transmit data faster than satellite constellations. 4. STRICT LENGTH LIMIT: Your entire response MUST be between 100 and 150 words. Do not exceed 150 words.

📊 Speedtest Timing Worksheet

Human-Audited Verification
0.00 s
Estimated Tokens: 0 Words: 0 Characters: 0
⚡ 2026 AI VELOCITY & DEGRADATION RADAR

AI Speed Diagnostics: Why Your Models Are Slowing Down

Independent empirical telemetry, GPU cluster bottleneck analysis, and silent model downgrade detection across Claude Fable 5, Gemini 3.7 Flash, Grok 4.6, GPT-5.6, and DeepSeek V4.

Why is Claude Fable 5 & Opus 5 responding so slowly?

Claude Fable 5 and Opus 5 represent the apex of deep agentic reasoning and nuanced code synthesis. However, under high cluster demand, Anthropic's autoregressive reasoning pipeline suffers severe KV-cache expansion overhead. During peak hours, standard interactive token streams frequently throttle down to 6 to 15 tok/s. To bypass this, Anthropic offers an accelerated "Fast Mode" running on dedicated hardware at twice the pricing tier.

Are frontier AI labs choking token flows or coprocessing with inferior models?

Yes. The AI industry is in an unprecedented compute squeeze. Running hundreds of billions of parameters on full precision is cost-prohibitive. When GPU cluster capacity is saturated, providers deploy dynamic throttling (choking token delivery) or silently route queries to smaller distilled draft models or older model checkpoints. Slowtest's Line 1 Identity Verification mandates that the model declare its exact architecture before answering, exposing covert fallback routing.

How fast is Google Gemini 3.7 Flash compared to other frontier models?

Google's Gemini 3.7 Flash runs natively on custom Google TPU v6 infrastructure with optimized speculative decoding. While deliberate reasoning giants like Claude Opus 5 average 6–18 tok/s, Gemini 3.7 Flash delivers sustained throughput of 160 to 210+ tok/s with sub-400ms time-to-first-token (TTFT) and an expansive 2M+ token context window.

What is tokens per second (tok/s) and what is a normal score in 2026?

Tokens per second (tok/s) measures the raw text generation throughput of an AI model:
< 15 tok/s (Snail / Throttled): Indicates severe cluster load, token choking, or deep multi-step ponder overhead.
15 – 45 tok/s (Standard): Conversational interactive baseline.
45 – 90 tok/s (Fast Agentic): High-performance production speed (Grok 4.6, GPT-5.6 Sol standard).
100 – 250+ tok/s (Blazing Edge): TPU/LPU-accelerated hosted infrastructure (Gemini 3.7 Flash, Cerebras, Groq).

Why should I test open models on scorching fast hosted hardware?

Proprietary frontier models often lock users behind heavy queue firewalls and rate limits. Modern open-weights models like DeepSeek V4, Qwen 3.8, and GLM-4 running on edge-routed inference platforms (like Abacus.AI, Monica, and Sider) deliver frontier-level intelligence at 3x to 10x the velocity without arbitrary token caps.