SkillsBench research · view source on GitHub ↗

Level 3 material is partial here

This page is rendered on-site from the skillsbench-history branch, but the underlying skill packages it discusses (seed/, best/, PROCESS.md under artifacts/) are not folded into the site yet — follow the GitHub ↗ links below to see them.

shock-analysis-demand — the same lost-fix pattern, with the mechanism proven

Status: DONE · Category: finance-economics / macroeconomic-analysis · Source run: c2 (run_task_shock-analysis-demand_v2)

Best: cand_0004 @ 0.900 · Δ vs seed: +0.900 · Held-out test: 0.900 · Iterations: 4

seed cand_0001 cand_0002 cand_0003 cand_0004 best
0.000 0.000 0.000 0.400 0.900 0.900

Numbers above are generated from results/results.json. Val rewards unless labelled test.

Material: - best/PROCESS.md — the optimizer's own write-up of the winning iteration: per-trial ground truth, ranked failure clusters with root causes, the kept edit, and what it deliberately skipped - best/ vs seed/ — diff these two trees to see exactly what changed - per-task-logs/shock-analysis-demand.md — per-trial reward vectors, from the earlier 43-task sweep (numbers may differ from above: different run) - evidence/shock-analysis-demand-optimizer-regression/ — full evidence bundle (journal, diffs, run reports)

This task has a full evidence bundle. The complete analysis — both runs' journals, the winning script, the accepted-candidate diffs, and the proof the seed never received the fix — is in evidence/shock-analysis-demand-optimizer-regression/. This report is the short version and the pointer; that bundle is the record.

Claim: cap-evolve found a validated fix worth +0.9 on this task, and an independent re-run with a weaker optimizer model scored +0.0 — despite naming the winning lever, correctly, four separate times, and having 50% more budget. The mechanism is proven rather than inferred: the re-run's seed skill demonstrably never received the prior run's accepted improvement.

The two runs

run worktree optimizer model iterations seed val best val best cand test Δ
prior c2, run_task_shock-analysis-demand_v2 claude-opus-4-8 4 0.0 0.9 cand_0004 +0.9
re-run c4, run_task_shock-analysis-demand_c4v1 claude-opus-4-6 6 0.0 0.0 seed (none accepted) +0.0

Both start from val 0.0 — that is not the regression, it is the same un-improved seed both times. The difference is entirely in what each run's optimizer achieved with its budget.

What changed

The winning run's fix was a two-step discovery, and the order matters:

  1. cand_0003 (0.0 -> 0.4) — switched levers from prose to code, shipping a new 261-line xlsx/scripts/build_shock_model.py that scaffolds the exact graded workbook structure: all 5 required sheets with correct column layout, every calculated cell written as a real Excel formula rather than a hardcoded value. The optimizer fetched the real verifier, template and oracle and confirmed the scaffolder scored 7/7 against the actual grading logic before proposing it.
  2. cand_0004 (0.4 -> 0.9) — diagnosed that reward was still flaky not because the script was wrong but because the agent often didn't run it: "the 4 passing trials ran it; all 6 failing trials did not." The fix was purely behavioural — a "START HERE" block making the scaffolder the mandatory first action — with no change to the script.

What worked

Verifying against the real grader before proposing. The scaffolder was checked 7/7 against the actual verifier, not against the optimizer's belief about the verifier. That is what made the 0.4 real rather than lucky.

Separating "the fix is wrong" from "the fix isn't being run." cand_0004 is the more instructive of the two: a correct artifact with a 40% adoption rate looks exactly like a broken artifact in the reward column. Distinguishing them required reading which trials invoked the script, and the remedy was a trigger change, not a code change.

What didn't

The re-run spent six iterations on levers the prior run had already refuted — prose, reference docs, and small helper/validator scripts (fetch_macro_data.py, inspect_xlsx.py, validate_formulas.py, validate_workbook.py). Every candidate scored Δ=+0.000. None of them was a structural scaffolder that writes the required sheets and formulas.

And it knew. Its own "focus next iteration" notes named the winning lever after four separate candidates — "a more complete workbook-scaffolding script that generates the full sheet structure", then again, then again, then "a script-heavy approach... might be needed despite overfitting risk." It identified the right move four times and never spent an iteration on it before the budget ran out.

What we can (or can't) learn

Can — cross-run knowledge is currently lost, and this is the proof, not the hypothesis. diffs/c4-seed-vs-c2-winning-skill.diff in the bundle shows the re-run's seed xlsx/SKILL.md lacks both the "START HERE" block and the macro-shock trigger clause, and the seed tree contains no build_shock_model.py at all. Accepted improvements are not merged back into the shared capability, so every run re-derives from zero. Same conclusion as flink-query, but here the absence is diffed rather than inferred.

Can — the optimizer can name the right lever without pulling it. Four deferrals of an explicitly-identified fix points at something specific and fixable: either lever-selection is too conservative about "script-heavy... despite overfitting risk", or budget allocation should weight later iterations toward a lever that has been deferred more than once.

Can — optimizer model strength is a real, isolatable variable. This is the cleanest such comparison on the branch: same task, same skill lineage, same failure cluster, more iterations for the weaker model, and the only deliberate difference is opus-4-8 -> opus-4-6.

Can't — call the re-run broken. It executed correctly. This is a legitimate, reproducible negative result, and filing it as a harness bug would have been wrong.

Can't — read +0.0 as "cap-evolve cannot do this task." It did do it, at +0.9, verified against the real grader.

Open experiment the bundle proposes: re-run this task with opus-4-6 while handing it the prior run's journal, to separate "can't discover the lever cold" from "can't execute it when told." Until that runs, the model-strength conclusion is one data point.