Last week a few of us pretrained a 15M-parameter language model from scratch — past its Chinchilla-optimal token budget — with no cluster, no server, and no budget. The training coordinator is a GitHub Actions cron job. The gradients are pull requests.
Source: [Hacker News](https://news.ycombinator.com/item?id=49141174)