SkillsBench research · view source on GitHub ↗
This page is rendered on-site from the skillsbench-history branch, but the underlying skill packages it discusses (seed/, best/, PROCESS.md under artifacts/) are not folded into the site yet — follow the GitHub ↗ links below to see them.
shock-analysis-demand — the same lost-fix pattern, with the mechanism proven
Status: DONE · Category: finance-economics / macroeconomic-analysis · Source run: c2 (run_task_shock-analysis-demand_v2)
Best: cand_0004 @ 0.900 · Δ vs seed: +0.900 · Held-out test: 0.900 · Iterations: 4
| seed | cand_0001 | cand_0002 | cand_0003 | cand_0004 | best |
|---|---|---|---|---|---|
| 0.000 | 0.000 | 0.000 | 0.400 | 0.900 | 0.900 |
Numbers above are generated from results/results.json. Val rewards unless labelled test.
Material:
- best/PROCESS.md — the optimizer's own write-up of the winning iteration: per-trial ground truth, ranked failure clusters with root causes, the kept edit, and what it deliberately skipped
- best/ vs seed/ — diff these two trees to see exactly what changed
- per-task-logs/shock-analysis-demand.md — per-trial reward vectors, from the earlier 43-task sweep (numbers may differ from above: different run)
- evidence/shock-analysis-demand-optimizer-regression/ — full evidence bundle (journal, diffs, run reports)
This task has a full evidence bundle. The complete analysis — both runs' journals, the winning script, the accepted-candidate diffs, and the proof the seed never received the fix — is in
evidence/shock-analysis-demand-optimizer-regression/. This report is the short version and the pointer; that bundle is the record.
Claim: cap-evolve found a validated fix worth +0.9 on this task, and an independent re-run with a weaker optimizer model scored +0.0 — despite naming the winning lever, correctly, four separate times, and having 50% more budget. The mechanism is proven rather than inferred: the re-run's seed skill demonstrably never received the prior run's accepted improvement.
The two runs
| run | worktree | optimizer model | iterations | seed val | best val | best cand | test Δ |
|---|---|---|---|---|---|---|---|
| prior | c2, run_task_shock-analysis-demand_v2 |
claude-opus-4-8 |
4 | 0.0 | 0.9 | cand_0004 |
+0.9 |
| re-run | c4, run_task_shock-analysis-demand_c4v1 |
claude-opus-4-6 |
6 | 0.0 | 0.0 | seed (none accepted) | +0.0 |
Both start from val 0.0 — that is not the regression, it is the same un-improved seed both times. The difference is entirely in what each run's optimizer achieved with its budget.
What changed
The winning run's fix was a two-step discovery, and the order matters:
cand_0003(0.0 -> 0.4) — switched levers from prose to code, shipping a new 261-linexlsx/scripts/build_shock_model.pythat scaffolds the exact graded workbook structure: all 5 required sheets with correct column layout, every calculated cell written as a real Excel formula rather than a hardcoded value. The optimizer fetched the real verifier, template and oracle and confirmed the scaffolder scored 7/7 against the actual grading logic before proposing it.cand_0004(0.4 -> 0.9) — diagnosed that reward was still flaky not because the script was wrong but because the agent often didn't run it: "the 4 passing trials ran it; all 6 failing trials did not." The fix was purely behavioural — a "START HERE" block making the scaffolder the mandatory first action — with no change to the script.
What worked
Verifying against the real grader before proposing. The scaffolder was checked 7/7 against the actual verifier, not against the optimizer's belief about the verifier. That is what made the 0.4 real rather than lucky.
Separating "the fix is wrong" from "the fix isn't being run." cand_0004 is the more
instructive of the two: a correct artifact with a 40% adoption rate looks exactly like a broken
artifact in the reward column. Distinguishing them required reading which trials invoked the
script, and the remedy was a trigger change, not a code change.
What didn't
The re-run spent six iterations on levers the prior run had already refuted — prose, reference
docs, and small helper/validator scripts (fetch_macro_data.py, inspect_xlsx.py,
validate_formulas.py, validate_workbook.py). Every candidate scored Δ=+0.000. None of them was
a structural scaffolder that writes the required sheets and formulas.
And it knew. Its own "focus next iteration" notes named the winning lever after four separate candidates — "a more complete workbook-scaffolding script that generates the full sheet structure", then again, then again, then "a script-heavy approach... might be needed despite overfitting risk." It identified the right move four times and never spent an iteration on it before the budget ran out.
What we can (or can't) learn
Can — cross-run knowledge is currently lost, and this is the proof, not the hypothesis.
diffs/c4-seed-vs-c2-winning-skill.diff in the bundle shows the re-run's seed xlsx/SKILL.md
lacks both the "START HERE" block and the macro-shock trigger clause, and the seed tree contains
no build_shock_model.py at all. Accepted improvements are not merged back into the shared
capability, so every run re-derives from zero. Same conclusion as
flink-query, but here the absence is diffed rather than inferred.
Can — the optimizer can name the right lever without pulling it. Four deferrals of an explicitly-identified fix points at something specific and fixable: either lever-selection is too conservative about "script-heavy... despite overfitting risk", or budget allocation should weight later iterations toward a lever that has been deferred more than once.
Can — optimizer model strength is a real, isolatable variable. This is the cleanest such
comparison on the branch: same task, same skill lineage, same failure cluster, more iterations
for the weaker model, and the only deliberate difference is opus-4-8 -> opus-4-6.
Can't — call the re-run broken. It executed correctly. This is a legitimate, reproducible negative result, and filing it as a harness bug would have been wrong.
Can't — read +0.0 as "cap-evolve cannot do this task." It did do it, at +0.9, verified against the real grader.
Open experiment the bundle proposes: re-run this task with opus-4-6 while handing it the
prior run's journal, to separate "can't discover the lever cold" from "can't execute it when
told." Until that runs, the model-strength conclusion is one data point.