Do you think we're trading off model size, and model capacity, for compute? Take looped transformers, which GPT-6 Astra is rumored to be based on. They increase compute while keeping the model's capacity almost fixed.
Source: [Hacker News](https://news.ycombinator.com/item?id=50003575)