SkillsBench task-by-task-87 — full-benchmark heatmap

87 tasks · dedicated LSF host per task (CCC) or dedicated Docker container (Mac) · 10 trials per candidate · up to 4-iteration budget · Sonnet-5 agent + Opus-4.8 optimizer. Data merged from four sources. Badges — source: c3v1 (CCC, bucket-1/2/3 unblocks, freshest) · c2v2 (CCC, 6 clean reruns) · c1v5 (CCC, post-cache-clear) · c1v1 (CCC, original 43) · c1tA (CCC, Track A) · mac-v2 (Boaz's Mac, 7 CCC-blocked retried); state: no-signal (task scored 0.0 for both baseline and every candidate — cap-evolve had no lever).

0.00
1.00

Click any task name to open its report — what changed, what worked, what didn't, and what we can (or can't) learn from that task. See reports/README.md for how far along that is: tasks already at 1.0 on seed, and the four no-signal tasks that are infrastructure failures, are answered in full; most of the rest are still placeholders.


c4 re-run — prev (c1/c2/c3) vs current (c4)

10 tasks re-run in a fresh worktree (intake_skillbench_c4) against a from-scratch seed capability, up to 6-iteration budget, optimizer model claude-opus-4-6 (down from claude-opus-4-8 on the prior run of these same tasks). "prev" is each task's own row from the full-87 heatmap above (same source badges apply); "curr" is the c4 run. cap-evolve's optimizer state does not carry forward between independent runs against a shared skill — a fresh run has to re-derive any fix from scratch within its budget, so a plateau here is lost progress, not a new bug. See evidence/shock-analysis-demand-optimizer-regression for a fully-worked example (shock-analysis-demand, Δ +0.9 → +0.0).

0.00
1.00