FP8 gave us a clean 1. 5x on Qwen3-8B serving throughput on an RTX PRO 6000 Blackwell (1,725 to 2,597 tok/s at concurrency 32, vLLM). The uncomfortable question is always the same: did the model get dumber.

Source: [Dev.to](https://dev.to/conatusai/did-fp8-make-the-model-dumber-a-per-prompt-regression-check-for-quantized-serving-595f)

Sponsored