Skip to main content
CodingAgentBench

CodingAgentBench · launching

The public site opens soon

You are looking at real published results — every number is a scored run, not a preview. The public site opens soon. We email you once when it lands. Nothing else.

One email when it lands

We send one email, ever, and never share the list. No newsletter, no drip, no upsell.

Live in the benchmark today

  • Ten coding-agent CLIs measured across fourteen open-weight models.
  • 3,465 real scored cells across integrity, mutations, and polyglot tasks.
  • An authentic terminal recording for every cell — watch each run replay.
  • Pinned container images with a sha256 digest on every row.
  • Leaderboard, Pareto, cube, and per-cell receipts over published results.

At public launch

  • Indexed public pages and a day-1 announcement.
  • Confidence intervals from k=3 reruns per cell, BCa bootstrap.
  • Held-out task seeds with a CVE-style embargo schedule.
  • Day-1 model launch numbers in lab partner blog posts.
  • A Zenodo DOI cut against the v0.1 methodology pin.