SkillsBench research · view source on GitHub ↗
This page is rendered on-site from the skillsbench-history branch, but the underlying skill packages it discusses (seed/, best/, PROCESS.md under artifacts/) are not folded into the site yet — follow the GitHub ↗ links below to see them.
flink-query — a real fix, found once and then lost on re-run (0.9 -> 0.2)
Status: DONE · Category: software-engineering / implementation · Source run: c3 (run_task_flink-query_v1_DONE)
Best: cand_0004 @ 0.900 · Δ vs seed: +0.900 · Held-out test: 0.900 · Iterations: 4
| seed | cand_0001 | cand_0002 | cand_0003 | cand_0004 | best |
|---|---|---|---|---|---|
| 0.000 | 0.100 | 0.100 | 0.500 | 0.900 | 0.900 |
Numbers above are generated from results/results.json. Val rewards unless labelled test.
Material:
- best/PROCESS.md — the optimizer's own write-up of the winning iteration: per-trial ground truth, ranked failure clusters with root causes, the kept edit, and what it deliberately skipped
- best/ vs seed/ — diff these two trees to see exactly what changed
Claim: cap-evolve found a precise, oracle-verified correctness fix on this task and drove it to 0.9. An independent re-run with a weaker optimizer model never rediscovered it, plateaued at 0.2, and spent all six iterations on the wrong lever. Optimizer state does not carry across runs — so a fix that is not merged back into the seed skill is a fix that can be lost.
The two runs
| run | source | seed | best val | best cand | held-out test | optimizer model |
|---|---|---|---|---|---|---|
| 87-task sweep | c3-v1 |
0.000 | 0.900 (cand_0004) | cand_0004 |
0.900 | claude-opus-4-8 |
| c4 re-run | c4v1 |
0.000 | 0.200 (cand_0001) | cand_0001 |
0.000 | claude-opus-4-6 |
Both against the same task, same deterministic 644-line oracle output. The recorded row in the auto block above is the 0.9 run.
What changed
The winning run climbed in two distinct stages, and only the second one was about correctness.
From best/PROCESS.md:
Stage 1 — get it to build (0.1 -> 0.5, cand_0003). Maven was rate-limiting with HTTP 429.
cand_0002 had shipped the mirror fix as a separate script and was rejected — agents never ran
it. cand_0003 moved the same fix inline into the skill body, added a "don't bypass the build
with javac" warning, and corrected the skill's own trigger description. Same fix, different
delivery, 5x the reward.
Stage 2 — fix the actual output bug (0.5 -> 0.9, cand_0004). The optimizer read the task's
oracle/solve.sh and verifier/test_outputs.py directly, then diffed a failing trial's real
output against the 644-line expected file. It found the agents were defining "job finished" as
any terminal event (FAIL || FINISH || KILL || LOST) instead of FINISH only — emitting 902
rows, 258 too many, none missing. It also found the skill text itself said "completion/terminal
event", actively steering agents into the bug.
The edit replaced that phrase with two explicit sub-rules: "finished" means the FINISH event
only, not the union of terminal states; and no fallback/Long.MAX_VALUE end-of-stream emit. It
verified every one of the five already-passing trials was already filtering to FINISH only —
so the edit narrowed exactly the wrong behaviour and could not disturb a working path.
What worked
Diagnosis against ground truth, not against intuition. The 0.5 -> 0.9 jump came from reading the oracle and comparing actual-vs-expected output, which turned a vague "output is wrong" into a one-line predicate error with a known row count. That is the single highest-leverage move in this task's history.
Fixing the skill's own misleading wording. The bug was partly caused by the skill saying "terminal event". The optimizer found the instruction that was producing the failure, not just the failure — a class of fix only available to something reading the skill and the trajectories together.
Inline over script, for a fix agents must actually apply. The same Maven fix was rejected as a
script and accepted inline. Consistent with shock-analysis-supply's finding in the opposite
direction (a gate works better as a script), the real rule is about whether the agent will
execute it on the path it actually takes.
What didn't
The re-run never got past infra triage. In c4v1, cand_0001 (0.2) correctly noted Maven 429
rate-limiting and made no edits. Then cand_0002 through cand_0006 — five consecutive
iterations — all pulled the same lever in variations: rewrite the skill's trigger description,
ship a Maven mirror script or inline commands, on the theory that the skill was never firing.
None of them touched the completion-join predicate. The optimizer never performed the failing-trace-vs-oracle comparison that unlocked 0.9 in the prior run, and its own final entry concludes the result is plausibly "pure variance" and that the skill may simply never trigger under the agent's judgement.
The test = 0.0 on the c4 row is not a real signal — it reflects Maven/podman infra failures
during the test phase, per that run's own journal. Do not read it as a regression to zero
capability.
What we can (or can't) learn
Can — the headline lesson: a fix that lives only in a run's candidate is not banked. cap-evolve's optimizer state does not carry forward between independent runs against a shared seed skill. A fresh run must re-derive any fix from scratch within its own budget. So a plateau on re-run is lost progress, not a new bug — and the operational implication is that accepted fixes need to be merged back into the seed capability, or they will be re-paid for (or not found at all) every time.
Can — optimizer model capability is load-bearing, not incidental. The only deliberate
difference between these runs was opus-4-8 -> opus-4-6, and the weaker model ran 6 iterations
without finding what the stronger one found in 4. The distinguishing behaviour was specific:
reading the oracle and diffing real output. That is a concrete, checkable capability difference,
not a vague "worked better".
Can — five iterations of one refuted lever is a diagnosable failure mode. The c4 run kept
re-trying trigger-description edits after they stopped paying. Compare shock-analysis-supply,
where the optimizer explicitly refused a third attempt at a refuted cluster. Whatever produced
that restraint did not fire here.
Can't — conclude the task is hard, or that 0.2 is its level. The 0.9 is real and oracle-verified. The 0.2 measures a run that never diagnosed the actual bug.
Can't — read the 0.9 -> 0.2 gap as run-to-run noise in the way shock-analysis-supply's spread
is. This is a traceable causal story with a named mechanism (unfound fix, weaker optimizer,
wrong lever repeated), confirmed against both runs' journals — not a wide sampling distribution.
The two tasks look similar in the heatmap and are not the same phenomenon.