In the last entry I got Gemma-4's 128-expert MoE running on an inf2. 24xlarge and signed off with a cliffhanger: fitting it on a 2-core box "needs fp4 — a separate expedition. " This is that expedition.
Source: [Dev.to](https://dev.to/aws-builders/squeezing-a-26b-moe-onto-the-cheapest-inferentia-box-from-a-649hr-24xlarge-to-a-076hr-3bhg)