Skip to main content
CodingAgentBench

OLAP cube · 3,465 cells

Slice the data any way you want

All 3,465 cells. Pick a TUI, a model, a task family. Get instant answers.

What this page is

  • What you'll find: every (CLI x model x task) cell scored on quality, latency, and cost.
  • Who it's for: people picking a CLI and model combo that wins on their workload.
  • How to switch: toggle to Nerds mode for the full pivot table, heatmap, and brush filters.

Top three winners

  1. opencode + GPT-OSS 120B wins 8/25 tasks at $0.0000 per task and 14.8s wall time (strongest on polyglot: 7).
  2. Pi + GPT-OSS 120B wins 5/25 tasks at $0.0000 per task and 12.6s wall time (strongest on polyglot: 4).
  3. Copilot + mistralai/mistral-small-4-119b-2603 wins 2/25 tasks at $0.0000 per task and 10.7s wall time (strongest on mutations: 1).

A combo wins a task when its composite score is the highest of any CLI on that task. Cost is a $1.0 / MTok proxy on tokens-per-correct (we call it the flat-rate roll-up); real per-token prices land via harness/tracing/cost.py.

Where the wins land

cli / model Qwen3.5 GPT-OSS 120B Llama 3.3 70B
Aider
Codex
Copilot
Crush
Goose
opencode
OpenHands
Pi
Plandex
Qwen-Code
clis
10
models
3
tasks
25
peak
8

Each cell is one (CLI, model) pair. Darker cells win more tasks. All 10 CLIs are shown; these are three representative models of the 14 in the full cube, which spans 25 tasks across three categories (polyglot, mutations, integrity). Switch to Nerds mode for every model.

Take me back home

OLAP cube · DuckDB-Wasm · 3,465 cells

Cube Explorer

Slice the benchmark cube on TUI × model × task × plugin-stack. Switch between table, heatmap, scatter, and parallel coordinates. The URL is the source of truth — every view is shareable.

v0.1 · 2026-06-20
Methodology v0.1 | Pinned to image digests as of 2026-06-20
3465 rows · 10 TUIs · 14 models · 3 categories