SkillsBench research · view source on GitHub ↗
This page is rendered on-site from the skillsbench-history branch, but the underlying skill packages it discusses (seed/, best/, PROCESS.md under artifacts/) are not folded into the site yet — follow the GitHub ↗ links below to see them.
shock-analysis-supply — three runs, three different answers (0.9 / 0.3 / 0.667)
Status: DONE · Category: finance-economics / macroeconomic-analysis · Source run: c2 (run_task_shock-analysis-supply_v2)
Best: cand_0004 @ 0.300 · Δ vs seed: +0.300 · Held-out test: 0.200 · Iterations: 4
| seed | cand_0001 | cand_0002 | cand_0003 | cand_0004 | best |
|---|---|---|---|---|---|
| 0.000 | 0.200 | 0.200 | 0.200 | 0.300 | 0.300 |
Numbers above are generated from results/results.json. Val rewards unless labelled test.
Material:
- best/PROCESS.md — the optimizer's own write-up of the winning iteration: per-trial ground truth, ranked failure clusters with root causes, the kept edit, and what it deliberately skipped
- best/ vs seed/ — diff these two trees to see exactly what changed
- per-task-logs/shock-analysis-supply.md — per-trial reward vectors, from the earlier 43-task sweep (numbers may differ from above: different run)
Claim: this task's reward is dominated by run-to-run variance, not by skill quality. Three
independent runs of the same task produced best-val of 0.9, 0.3 and 0.667. The
canonical results.json row records the worst of the three. Any per-task conclusion drawn
from a single run of this task is unsafe.
The three runs
| run | source | seed | best val | best cand | held-out test | recorded where |
|---|---|---|---|---|---|---|
| 43-task sweep (batch 3) | c1 |
0.100 | 0.900 (cand_0003) | cand_0003 |
not run | per-task-logs/shock-analysis-supply.md |
| 87-task sweep | c2-v2 |
0.000 | 0.300 (cand_0004) | cand_0004 |
0.200 | results/results.json — the auto block above |
c4 re-run (6-iter budget, optimizer opus-4-6) |
c4v1 |
0.000 | 0.667 (cand_0003) | cand_0003 |
0.000 | DATA_C4 in ui/heatmap.html |
The reward is binary per trial (a trial scores 1.0 only if all 9 verifier tests pass), so with 10 trials each of these means is a count of passing trials out of 10 — 9/10, 3/10, and ~7/10 (6 of 9 evaluated). The spread is not a measurement artifact of a noisy metric; it is genuinely different numbers of trials succeeding.
What changed
The accepted fix in the recorded (c2-v2) run was a script, not prose — the pattern
cap-evolve is supposed to produce. From
best/PROCESS.md:
- New
xlsx/scripts/check_econ_model.py— a units/magnitude gate the agent runs after recalc. It locates the capital (K) and output (Real GDP) columns by header label, computes the capital-output ratio K/Y, and fails with the exact remedy when K/Y falls outside the economically plausible ~1–30 band. A missing×1000scaling yields K/Y ≈ 3000, so the gate catches precisely the observed bug. xlsx/SKILL.md— a workflow step that runs the gate, plus the K/Y invariant stated numerically.xlsx/references/economic_models.md§3 — the same K/Y check as the most reliable magnitude test.
What worked
The root cause was correctly identified and was behavioural, not a knowledge gap. Two
near-miss trials (t3, t4) failed only test_value_magnitudes: the agent linked Real GDP into
the production sheet without the ×1000 scaling, leaving GDP in billions (~80) while capital
was in millions (~300k). The TFP residual absorbed the offset so every formula test still
passed — only the final magnitudes were wrong, by a factor of 1000.
Critically, the optimizer verified that both agents had read the reference rule requiring
×1000 and skipped it anyway. That is what justified shipping a script gate rather than more
prose: an instruction that is read and ignored does not get fixed by rewording it. The gate was
verified by running it against the oracle workbook (passes, K/Y 3.45) and a synthesized broken
copy (fails, naming the exact remedy), and confirmed to no-op on the empty template and on the
sibling shock-analysis-demand model.
What didn't
The catastrophic cluster was never solved, and two levers were spent proving prose can't fix
it. Five of ten trials (t0, t1, t6, t8, t9) never reached a build at all — budget exhaustion
during data ingestion, an agent entering plan mode and stalling on an unavailable
AskUserQuestion, and an IMF HTTP-403 loop. Against that cluster:
cand_0002shipped a data-fetch script with concrete SDMX/ECB/PWT endpoints and a closed-form HP filter → rejected, changed nothing. The endpoints were not the bottleneck.cand_0003shipped sequencing prose ("build a skeleton first, save early, don't enter plan mode, finish autonomously") → rejected, reverted. Prose did not move those trials.
The optimizer then deliberately declined to attempt cluster 2 again, on the grounds that both plausible levers were already refuted and a speculative third would risk the one verified win. That restraint is correct behaviour and worth noting: it chose not to manufacture a change.
On the c4 run's apparent "decay" after 0.667
In the c4 re-run the candidate sequence reads 0.0, 0.0, 0.667, 0.4, 0.0, 0.0 — which looks
like the fix degrading. It is not. cand_0003 (0.667) was accepted and became champion; the
three candidates after it were each independently rejected and reverted to that champion.
Their own scores are a failed follow-on experiment and trial noise, not a decline off 0.667.
The Best column never regresses, and that is the column that means anything.
What we can (or can't) learn
Can: a read-but-ignored instruction is a behavioural failure, and the fix that works is an executable gate that fails loudly, not better wording. This task is a clean instance of that lesson, verified rather than asserted.
Can: the optimizer correctly refused to spend a fourth iteration on a cluster whose levers were refuted — evidence the accept/reject gate suppresses busywork rather than rewarding churn.
Can't: anything about this task's absolute difficulty or about cap-evolve's expected lift on it. With best-val of 0.9, 0.3 and 0.667 across three runs, the recorded 0.3 is not representative, and the +0.300 delta in the auto block above should not be quoted as the result for this task. It is one draw from a wide distribution, and it happens to be the lowest.
Can't: treat the 0.9 from the 43-task sweep as directly comparable either — it ran under a different seed capability and batch configuration. The honest statement is that this task has never been run enough times, under one fixed configuration, to have a defensible number.
Methodological consequence worth carrying upward: the 87-task aggregate takes one recorded run per task. Where a task's run-to-run spread is this wide, the aggregate is a mix of draws rather than a single experiment. That does not invalidate the headline, but it does mean per-task deltas are much softer than they look, and it is a further reason the +2.5 pp vs EvoSkill claim was withdrawn.