AI & ML
Two "Codex CLI" models on the same benchmark: the harness hides the model
Cole Halton Dev.to (EN Zone)
2 views
Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name.
Same harness, two different brains:
GPT-6 Astra on Codex CLI: 33.8% resolution
GPT-5.6 Sol on Codex CLI: 16.2% resolution
Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Neither reading is right, because the harness is just the routing layer, not the thing that does the thinking.
It's the same on the Claude side. Fable 5.1 on Claude Code is 38.8%. GLM 5.3 on Claude Code is 28.8%. One harness name, ten points apart. Real-SWE is careful about this, actually: they frame every result as a model-and-harness combination, not a model in isolation. That's the honest way to present it, and most vendors won't do it because a low number under their own tooling looks bad.
This matters for teams actually shopping for coding agents, because "we use Claude Code" says nothing about the skill of the agent you get. You've picked a route, not a brain. The model swap is the biggest lever, and it's completely invisible in the marketing.
The eval habit I'd take from this: when someone hands you a benchmark score, pin both halves. Model. Harness. Conventions, context carry-over, tool loop, judge. If a tool vendor won't tell you which model a number belongs to, that's a red flag, not a detail. The scaffold can move a score more than reasoning effort does.
Read original: https://dev.to/cole_halton_42f71d71b809b/two-codex-cli-models-on-the-same-benchmark-the-harness-hides-the-model-1jn
← Previous
Giving a coding agent more time barely helps
Next →
Cutting PR review time is really changing where review happens
Related
From Claude Project to Hybrid AI Agent: Lessons from a Real-World Content Workflow
AI & ML
1
DEV Community
Nvidia จ่อทุ่ม 1 หมื่นล้านเข้า IPO Anthropic คำถามคือใครเป็นลูกค้าของใคร
AI & ML
2
DEV Community
นักคณิตศาสตร์เหรียญ Fields 25 คนบอกว่า AI กำลังทำร้ายคณิตศาสตร์
AI & ML
1
DEV Community
Chasing Quantum States Through Time: A Tour of Time-Evolution Methods in TensorCircuit-NG
AI & ML
1
DEV Community
Comments0
No comments yet — be the first