SkillsBench research · view source on GitHub ↗

Level 3 material is partial here

This page is rendered on-site from the skillsbench-history branch, but the underlying skill packages it discusses (seed/, best/, PROCESS.md under artifacts/) are not folded into the site yet — follow the GitHub ↗ links below to see them.

shock-analysis-supply — three runs, three different answers (0.9 / 0.3 / 0.667)

Status: DONE · Category: finance-economics / macroeconomic-analysis · Source run: c2 (run_task_shock-analysis-supply_v2)

Best: cand_0004 @ 0.300 · Δ vs seed: +0.300 · Held-out test: 0.200 · Iterations: 4

seed cand_0001 cand_0002 cand_0003 cand_0004 best
0.000 0.200 0.200 0.200 0.300 0.300

Numbers above are generated from results/results.json. Val rewards unless labelled test.

Material: - best/PROCESS.md — the optimizer's own write-up of the winning iteration: per-trial ground truth, ranked failure clusters with root causes, the kept edit, and what it deliberately skipped - best/ vs seed/ — diff these two trees to see exactly what changed - per-task-logs/shock-analysis-supply.md — per-trial reward vectors, from the earlier 43-task sweep (numbers may differ from above: different run)

Claim: this task's reward is dominated by run-to-run variance, not by skill quality. Three independent runs of the same task produced best-val of 0.9, 0.3 and 0.667. The canonical results.json row records the worst of the three. Any per-task conclusion drawn from a single run of this task is unsafe.

The three runs

run source seed best val best cand held-out test recorded where
43-task sweep (batch 3) c1 0.100 0.900 (cand_0003) cand_0003 not run per-task-logs/shock-analysis-supply.md
87-task sweep c2-v2 0.000 0.300 (cand_0004) cand_0004 0.200 results/results.json — the auto block above
c4 re-run (6-iter budget, optimizer opus-4-6) c4v1 0.000 0.667 (cand_0003) cand_0003 0.000 DATA_C4 in ui/heatmap.html

The reward is binary per trial (a trial scores 1.0 only if all 9 verifier tests pass), so with 10 trials each of these means is a count of passing trials out of 10 — 9/10, 3/10, and ~7/10 (6 of 9 evaluated). The spread is not a measurement artifact of a noisy metric; it is genuinely different numbers of trials succeeding.

What changed

The accepted fix in the recorded (c2-v2) run was a script, not prose — the pattern cap-evolve is supposed to produce. From best/PROCESS.md:

What worked

The root cause was correctly identified and was behavioural, not a knowledge gap. Two near-miss trials (t3, t4) failed only test_value_magnitudes: the agent linked Real GDP into the production sheet without the ×1000 scaling, leaving GDP in billions (~80) while capital was in millions (~300k). The TFP residual absorbed the offset so every formula test still passed — only the final magnitudes were wrong, by a factor of 1000.

Critically, the optimizer verified that both agents had read the reference rule requiring ×1000 and skipped it anyway. That is what justified shipping a script gate rather than more prose: an instruction that is read and ignored does not get fixed by rewording it. The gate was verified by running it against the oracle workbook (passes, K/Y 3.45) and a synthesized broken copy (fails, naming the exact remedy), and confirmed to no-op on the empty template and on the sibling shock-analysis-demand model.

What didn't

The catastrophic cluster was never solved, and two levers were spent proving prose can't fix it. Five of ten trials (t0, t1, t6, t8, t9) never reached a build at all — budget exhaustion during data ingestion, an agent entering plan mode and stalling on an unavailable AskUserQuestion, and an IMF HTTP-403 loop. Against that cluster:

The optimizer then deliberately declined to attempt cluster 2 again, on the grounds that both plausible levers were already refuted and a speculative third would risk the one verified win. That restraint is correct behaviour and worth noting: it chose not to manufacture a change.

On the c4 run's apparent "decay" after 0.667

In the c4 re-run the candidate sequence reads 0.0, 0.0, 0.667, 0.4, 0.0, 0.0 — which looks like the fix degrading. It is not. cand_0003 (0.667) was accepted and became champion; the three candidates after it were each independently rejected and reverted to that champion. Their own scores are a failed follow-on experiment and trial noise, not a decline off 0.667. The Best column never regresses, and that is the column that means anything.

What we can (or can't) learn

Can: a read-but-ignored instruction is a behavioural failure, and the fix that works is an executable gate that fails loudly, not better wording. This task is a clean instance of that lesson, verified rather than asserted.

Can: the optimizer correctly refused to spend a fourth iteration on a cluster whose levers were refuted — evidence the accept/reject gate suppresses busywork rather than rewarding churn.

Can't: anything about this task's absolute difficulty or about cap-evolve's expected lift on it. With best-val of 0.9, 0.3 and 0.667 across three runs, the recorded 0.3 is not representative, and the +0.300 delta in the auto block above should not be quoted as the result for this task. It is one draw from a wide distribution, and it happens to be the lowest.

Can't: treat the 0.9 from the 43-task sweep as directly comparable either — it ran under a different seed capability and batch configuration. The honest statement is that this task has never been run enough times, under one fixed configuration, to have a defensible number.

Methodological consequence worth carrying upward: the 87-task aggregate takes one recorded run per task. Where a task's run-to-run spread is this wide, the aggregate is a mix of draws rather than a single experiment. That does not invalidate the headline, but it does mean per-task deltas are much softer than they look, and it is a further reason the +2.5 pp vs EvoSkill claim was withdrawn.