An open model edged out a closed one on coding tasks — and it barely matters
The latest leaderboard reshuffled an open model past a proprietary one. What is behind the number, and why the number is misleading.
A fresh report put an open model at the top of a coding benchmark, ahead of a proprietary one by roughly two percentage points. The news spread in a couple of hours, but it deserves a closer read.
Where those two points came from
Gaps that size almost always sit inside the measurement noise. Coding benchmarks are graded by automated tests, and automated tests are fragile: change the wording of a task or the order of examples and the ranking flips. The authors say so honestly in a footnote on page eight — a page almost nobody reaches.
What genuinely changed
What matters is not the leaderboard slot but the fact that the weights now fit on a single machine. A year ago this level of quality meant a server rack and a cloud contract.
- The weights fit in the memory of an ordinary workstation.
- Throughput is good enough for interactive work.
- The licence permits commercial use.
The catch
Public-benchmark quality and real-world quality diverge sharply. Public tasks are labelled and cleanly phrased; real requests arrive misspelled, tangled with context, and without a stated goal. There the gap between models is usually wider than any leaderboard suggests — in both directions.
What to look at instead of leaderboards
Look at latency, price per million tokens, and behaviour on your own work. Take a hundred real requests from your week, run them through both models, and compare with your own eyes. That costs an evening and tells you more than any ranking.