SkillsBench research · view source on GitHub ↗

Zero-shot skill transfer — 8-fold pilot (transfer_eval_v1)

Generated 2026-09-02 · 8 folds · 5 distinct tasks · dedicated LSF host per fold (CCC) · cap-evolve run --max-iterations 0 (pure evaluation, no optimizer loop) · Sonnet-5 agent.

Status: 8 of 8 folds complete.

What the experiment asks

Take the frozen winning skill that cap-evolve produced for task A, hand it to a different task B, and evaluate it zero-shot (--max-iterations 0, no adaptation). Does a skill optimized for one task carry over to another?

The comparison of interest is the transferred skill's score on B versus B's own native seed — i.e. what B scored before any optimization at all. If the donor skill beats B's native seed, transfer helped; if it scores below, the donor skill is actively worse than starting from scratch on B.

Results

transfer reward = the donor skill's held-out test_reward on the test task. native seed / native optimized = the test task's own numbers from results/results.json (the task-by-task-87 run).

# skill from (train task) evaluate on (test task) job status transfer reward native seed native optimized transfer − native seed
1 shock-analysis-demand shock-analysis-supply 555427 DONE 0.1 0.0 0.2 +0.1
2 shock-analysis-demand weighted-gdp-calc 540431 DONE 0.2 0.8 1.0 ‡ −0.6
3 shock-analysis-supply shock-analysis-demand 543167 DONE 0.0 0.0 0.9 0.0
4 shock-analysis-supply weighted-gdp-calc 540433 DONE 0.3 0.8 1.0 ‡ −0.5
5 weighted-gdp-calc shock-analysis-demand 543168 DONE 0.0 0.0 0.9 0.0
6 weighted-gdp-calc shock-analysis-supply 543169 DONE 0.1 0.0 0.2 +0.1
7 exam-block-sequencing paratransit-routing 543170 DONE 0.2 0.0 1.0 ‡ +0.2
8 paratransit-routing exam-block-sequencing 543171 DONE 0.4 0.1 1.0 ‡ +0.3

‡ In-loop val score, not a held-out test score. These four tasks are KILLED_ceiling in results.json with final_test: null — optimization saturated and the run was stopped before a held-out test eval ran. Rows 3 and 5 use held-out final_test (0.9); row 1/6's 0.2 is shock-analysis-supply's final_test (its val best was 0.3).

Reading of all 8 folds

Methodological caveat — ignore cap-evolve's own test_delta

Every fold reports best_id: "seed", test_reward == test_baseline_reward, and test_delta: 0.0. This is an artifact of --max-iterations 0, not a finding. With no optimizer iterations, the transfer project's "seed" is the frozen donor skill, and it is the only artifact evaluated — so cap-evolve compares it against itself and the delta is trivially zero by construction. Any real transfer effect must be computed against results.json's native-seed column, as done above.

Per-fold raw results

Machine-readable: transfer_eval_8fold.json.

Per-fold logs live in the intake_skillbench_c5 worktree at results/transfer_eval_v1/<jobid>/cap-evolve.log, with cap-evolve run dirs at .capevolve/run_transfer_<train>_to_<test>_v{1,2}/.

Operational notes (CCC / LSF)

Worth recording because it cost several reruns:

Provenance