Hi guys, A few months ago, I've been training several simple models as a side hobby. I tried training N models at once on a single GPU, but OOM spikes kept crashing my runs. So I built my own simple way to manage this kind of training.
Source: [Hacker News](https://github.com/ceruleane/gpusched)