Hey HN! We all know that most public evals are saturated and are hard to trust. What matters is whether a model reliably works in your codebase in your actual day-to-day-work.

Source: [Hacker News](https://github.com/mupt-ai/self-bench)

Sponsored