Same model. Same prompt. Ten minutes apart.
Source: [Dev.to](https://dev.to/ptokito/the-third-price-i-measured-prompt-caching-across-393-llms-and-found-a-90-discount-hiding-behind-eck)
2 points, 0 comments on Hacker News
FP8 gave us a clean 1. 5x on Qwen3-8B serving throughput on an RTX PRO 6000 Blackwell (1,725 to 2,597 tok/s at concurrency 32, vLLM). The uncomfortable question is always the same: did the model get dumber.
1 points, 0 comments on Hacker News
Here is how I used to change a prompt, and I suspect it's how you do it too. Take a handful of real inputs — six felt like plenty, you can hold six in your head. Run the old prompt and the new prompt over them.
1 points, 0 comments on Hacker News
"Meat proxy" refers to someone who sends outputs straight from Claude or ChatGPT without analyzing them. It's caught on with techies on X.