Benchmark runs

Every ci/benchmarks execution (tau2 · swebench · skillsbench · rfe-creator), on PRs and manual runs. Each row is one benchmark in one run, measured against that run's own freshly-computed baseline (all of a tier's tasks are optimized together, in one run — nothing is pre-frozen or reused across runs); the Type column marks the run size — smoke and full are the standing ones, plus any ad-hoc tier a run declares — and Experiment marks which named study a run belongs to, where applicable. The Benchmark/Type/Experiment filters are built from the loaded data, and Type/Experiment list only the values that exist for the selected benchmark. Click a row to expand its per-task reward and per-iteration cost/latency detail.

See the hand-verified canonical numbers on the Results page; this is the raw CI log. Failed/cancelled runs are shown too, so infra breakage is visible. SkillsBench rows with a Report ↗ link in the UI column drill down to a rendered heatmap and (where it exists) per-task detail — start from the SkillsBench experiments hub.

Date Source Bench Type Experiment Algorithm Iters Trials Reward base→opt Eval $ Optimizer $ Latency Agent model Optimizer model Result UI