Benchmark runs
Every ci/benchmarks execution (tau2 · swebench · skillsbench · rfe-creator), on PRs and manual
runs. Each row is one benchmark in one run, measured against that run's own
freshly-computed baseline (all of a tier's tasks are optimized together, in one run —
nothing is pre-frozen or reused across runs); the Type column marks the
run size — smoke and full are the standing ones, plus any
ad-hoc tier a run declares — and Experiment marks which named study a
run belongs to, where applicable. The Benchmark/Type/Experiment filters are built from
the loaded data, and Type/Experiment list only the values that exist for the selected
benchmark. Click a row to expand its per-task reward and per-iteration cost/latency
detail.
See the hand-verified canonical numbers on the Results page; this is the raw CI log. Failed/cancelled runs are shown too, so infra breakage is visible. SkillsBench rows with a Report ↗ link in the UI column drill down to a rendered heatmap and (where it exists) per-task detail — start from the SkillsBench experiments hub.
Running now / queued
| Date | Source | Bench | Type | Experiment | Algorithm | Iters | Trials | Reward base→opt | Eval $ | Optimizer $ | Latency | Agent model | Optimizer model | Result | UI |
|---|
No runs recorded yet — trigger the suite (add a
benchmark-smoke-<bench> label to a PR, or Actions → Benchmarks).
Couldn't load benchmark history.