How scalably can we cheaply fine-tune small models for well-defined tasks? Frontier models are bad at clock reading. I fine-tuned a tiny 450M model from Liquid AI to match GPT-5.
Source: [Hacker News](https://huggingface.co/jadidbourbaki/time-wizard)