Built this cuz no library gave LLM API calls in prod a consistent way to retry, timeout, circuit-break and cache across multiple providers.
Source: [Hacker News](https://vernllm.vercel.app)
1 points, 0 comments on Hacker News
1 points, 0 comments on Hacker News
"Microsoft listened to my feedback!
1 points, 0 comments on Hacker News
1 points, 1 comments on Hacker News
I run an automated video pipeline that generates images locally, on my Mac. A 16GB M4. The image model is FLUX, a 12B-parameter thing that keeps ~7GB of weights resident in memory even quantized down to 4bit.