I saw some great results from really tiny LLM's, and ended up making a little benchmark so I could test them in the browser. Chrome, Firefox, Safari and Edge are all supported via web-GPU. I got very fast speeds 30-40 tok/sec even on limited hardware.
Source: [Hacker News](https://github.com/robss2020/microllm-lab)