I wanted to try using the larger models on my computer (32GB RAM, RTX 5080, Gen5 NVMe), but the best I could do was around 30B. So I started with the idea that it might be possible by taking advantage of the fact that MoE models use only some of the experts rather than all of them.
Source: [Hacker News](https://github.com/tmxkzm1925-max/MoE-Direct)