SkillsBench experiments

Three experiments cap-evolve has run against BenchFlow's SkillsBench (87 tasks, 8 categories). Each card is level 1 — a row on the Benchmarks dashboard — with links onward to level 2 (a heatmap or per-task table) and, where it exists, level 3 (a per-task report).

Level 3 is not fully on-site yet

Per-task reports below are rendered from skillsbench-history, but the skill packages they discuss (seed/, best/, PROCESS.md under artifacts/ — about 27 MB) are not folded into the site. Every report page links out to GitHub for that material; each such link is marked GitHub ↗.

task-by-task — one skill per task (87 tasks)

EvoSkills-style: each task gets its own optimized skill, seeded from the task's own environment/skills/. The full 3-level drill-down exists for this experiment.

Level 2 — summary

Aggregate pass rate, per-category breakdown, comparison to the EvoSkills 71.1% headline.

Level 2 — heatmap

Interactive per-task × per-candidate score grid, 87 tasks. Click a task name for its report.

Level 3 — per-task reports

87 reports: why a task moved (or didn't), and what it teaches.

transfer-eval-8fold — zero-shot skill transfer (8 folds)

Does a skill optimized for one task carry over to a different task, evaluated zero-shot (--max-iterations 0, no adaptation)? No level 3 — this experiment doesn't produce per-task skill artifacts of its own, only a comparison table.

Level 2 — summary

The full 8-fold writeup, reading, methodological caveats, and CCC/LSF operational notes.

Level 2 — heatmap

Transfer reward vs. each test task's own native seed/optimized score, 8 folds.

runnable-subset-43 — whole-suite optimization (43 tasks)

One shared skill package optimized across all 43 tasks that actually run on our CCC podman setup. Cancelled at iteration 4 (optimizer hang); best accepted candidate is cand_0002.

Level 2 — summary

Aggregate reward, iteration table, passing/partial/regressed task lists, cost and wall clock.

Level 2 — heatmap

Seed-baseline reward per task, annotated with cand_0002's pass/partial outcome. No full candidate × task matrix exists for this run — see the page for why.

More

Skills ↔ tasks map

All 87 tasks and their skills, numbered and cross-referenced, plus a per-category train/test split for skill optimization.

Task inventory

All 87 tasks by category, difficulty, task type, and shipped skills.