Verdict: Inception Labs' Mercury 2. 5 is the fastest LLM you can call from an API in September 2026 — 1,107 tokens per second on widely available NVIDIA GPUs, at a list price of $0. 20/$0.

Source: [Dev.to](https://dev.to/shaam_ai/fastest-llm-2026-mercury-25-beats-luna-and-haiku-2nef)

Sponsored