On an H100, tokens per watt drops 12x between 4K and 64K context. Agents live at the fat end of that curve. The fix comes from semiconductor architecture.
Source: [HackerNoon](https://hackernoon.com/tokens-per-watt-why-your-context-window-is-a-power-decision?source=rss)