Benchmark runs

Every ci/benchmarks execution (tau2 · swebench · skillsbench), on PRs and manual runs. Each row is one benchmark in one run, measured against that run's own freshly-computed baseline (all of a tier's tasks are optimized together, in one run — nothing is pre-frozen or reused across runs); the Type column marks the tier (smoke or full). Click a row to expand its per-task reward and per-iteration cost/latency detail.

See the hand-verified canonical numbers on the Results page; this is the raw CI log. Failed/cancelled runs are shown too, so infra breakage is visible.

Date Source Bench Type Iters Trials Reward base→opt Eval $ Optimizer $ Latency Agent model Optimizer model Result UI