PicoLM is an LLM inference engine written in C99. It currently supports llama-2, GPT-2, Qwen 3. 6/3.

Source: [Hacker News](https://github.com/whoreson/picolm/)

Sponsored