I keep seeing the same local LLM sizing mistake: "The model file is smaller than my GPU, so it should fit. " That is only the first check. A 24 GB GPU does not give your model a clean 24 GB memory budget.
Source: [Dev.to](https://dev.to/deep_mehta_b12a764b0ec6b5/why-a-24-gb-gpu-does-not-give-your-local-llm-24-gb-4b0k)