No synthetic puzzles, no contaminated suites: 25 tasks harvested from PRs merged in 2026 across five languages, graded by each project's own held-out tests. Four agents at stock settings. octomind with an open model solved 24/25 - ahead of Claude Code with Opus - while the same model in another...

Source: [Dev.to](https://dev.to/donk8r/benchmarking-ai-coding-agents-on-real-pull-requests-22k9)

Sponsored