87 tasks · dedicated LSF host per task (CCC) or dedicated Docker container (Mac) · 10 trials per candidate · up to 4-iteration budget · Sonnet-5 agent + Opus-4.8 optimizer. Data merged from four sources. Badges — source: c3v1 (CCC, bucket-1/2/3 unblocks, freshest) · c2v2 (CCC, 6 clean reruns) · c1v5 (CCC, post-cache-clear) · c1v1 (CCC, original 43) · c1tA (CCC, Track A) · mac-v2 (Boaz's Mac, 7 CCC-blocked retried); state: no-signal (task scored 0.0 for both baseline and every candidate — cap-evolve had no lever).
Click any task name to open its report — what changed, what worked, what didn't, and what we can (or can't) learn from that task. See reports/README.md for how far along that is: tasks already at 1.0 on seed, and the four no-signal tasks that are infrastructure failures, are answered in full; most of the rest are still placeholders.
10 tasks re-run in a fresh worktree (intake_skillbench_c4) against a from-scratch seed capability,
up to 6-iteration budget, optimizer model claude-opus-4-6 (down from
claude-opus-4-8 on the prior run of these same tasks). "prev" is each task's own row from the
full-87 heatmap above (same source badges apply); "curr" is the c4 run.
cap-evolve's optimizer state does not carry forward between independent runs against a shared skill —
a fresh run has to re-derive any fix from scratch within its budget, so a plateau here is lost progress,
not a new bug. See
evidence/shock-analysis-demand-optimizer-regression
for a fully-worked example (shock-analysis-demand, Δ +0.9 → +0.0).