SkillsBench experiments
Three experiments cap-evolve has run against BenchFlow's SkillsBench (87 tasks, 8 categories). Each card is level 1 — a row on the Benchmarks dashboard — with links onward to level 2 (a heatmap or per-task table) and, where it exists, level 3 (a per-task report).
Per-task reports below are rendered from skillsbench-history, but the skill
packages they discuss (seed/, best/, PROCESS.md under
artifacts/ — about 27 MB) are not folded into the site. Every report page links
out to GitHub for that material; each such link is marked GitHub ↗.
task-by-task — one skill per task (87 tasks)
EvoSkills-style: each task gets its own optimized skill, seeded from the task's own
environment/skills/. The full 3-level drill-down exists for this experiment.
Aggregate pass rate, per-category breakdown, comparison to the EvoSkills 71.1% headline.
Interactive per-task × per-candidate score grid, 87 tasks. Click a task name for its report.
87 reports: why a task moved (or didn't), and what it teaches.
transfer-eval-8fold — zero-shot skill transfer (8 folds)
Does a skill optimized for one task carry over to a different task, evaluated zero-shot
(--max-iterations 0, no adaptation)? No level 3 — this experiment doesn't
produce per-task skill artifacts of its own, only a comparison table.
The full 8-fold writeup, reading, methodological caveats, and CCC/LSF operational notes.
Transfer reward vs. each test task's own native seed/optimized score, 8 folds.
runnable-subset-43 — whole-suite optimization (43 tasks)
One shared skill package optimized across all 43 tasks that actually run on our CCC podman
setup. Cancelled at iteration 4 (optimizer hang); best accepted candidate is cand_0002.
Aggregate reward, iteration table, passing/partial/regressed task lists, cost and wall clock.
Seed-baseline reward per task, annotated with cand_0002's pass/partial outcome. No full candidate × task matrix exists for this run — see the page for why.