There is a problem with most benchmarks for coding agents: they are not software development. They are useful, of course. Give an agent an issue, run a test suite, check whether the patch passes.

Source: [Dev.to](https://dev.to/jjackbauer/800-million-tokens-for-36-benchmarking-dsh-deepseek-v4-pro-on-a-real-codebase-2a0o)

Sponsored