Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33. 8% resolution GPT-5.

Source: [Dev.to](https://dev.to/cole_halton_42f71d71b809b/two-codex-cli-models-on-the-same-benchmark-the-harness-hides-the-model-1jn)

Sponsored