Benchmark runs
Every ci/benchmarks execution (tau2 · swebench · skillsbench), on PRs and manual
runs. Each row is one benchmark in one run, measured against that run's own
freshly-computed baseline (all of a tier's tasks are optimized together, in one run —
nothing is pre-frozen or reused across runs); the Type
column marks the tier (smoke or full). Click a row to expand its
per-task reward and per-iteration cost/latency detail.
See the hand-verified canonical numbers on the Results page; this is the raw CI log. Failed/cancelled runs are shown too, so infra breakage is visible.
Running now / queued
| Date | Source | Bench | Type | Iters | Trials | Reward base→opt | Eval $ | Optimizer $ | Latency | Agent model | Optimizer model | Result | UI |
|---|
No runs recorded yet — trigger the suite (add a
benchmark-smoke label to a PR, or Actions → Benchmarks).
Couldn't load benchmark history.