wiki / concepts / skill-treatment-effect
Skill Treatment Effect
Machine ingest — raw context
loading…
~… tokensappend .md to any wiki URL for this view
Skill Treatment Effect
The Skill Treatment Effect refers to the measured marginal utility of injecting procedural knowledge packages (evals skills, SKILL.md) into an agent’s context during inference. Empirical benchmarks demonstrate that skills are not a universal productivity booster; their effectiveness depends heavily on domain specificity, task fit, and version compatibility.
flowchart LR
A[Raw Skill Injection] --> B{Task & Domain Fit}
B -->|Domain-Specific / Tool Procedure| C[+15% to +30% Lift]
B -->|Generic Software Engineering| D[+1.2% to +4.5% Marginal Lift]
B -->|Version Mismatch / Contradiction| E[-5% to -10% Regression & Bloat]
Empirical Evidence
SkillsBench (Cross-Domain Benchmark)
- Scope: 87 tasks across 8 domains evaluated over 18 model-harness configurations.
- Aggregate Result: Curated skills increased the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain).
- Domain Variance: Broad non-coding tasks saw substantial lift, but software engineering tasks experienced only a modest +4.5 pp improvement.
- Composition Law: Focused skills with $\le 3$ modules outperformed exhaustive or kitchen-sink bundles. Self-generated skills yielded zero net benefit.
SWE-Skills-Bench (Software Engineering Focus)
- Scope: 49 public SWE skills across ~565 task instances in real-world GitHub repositories with deterministic verifiers.
- Aggregate Result: Mean pass-rate gain across all skills was only +1.2%.
- Distribution:
- 39 of 49 skills: Zero pass-rate improvement while increasing token overhead by up to 451%.
- 7 specialized skills: Delivered meaningful gains (up to +30%) by providing toolchain-specific API knowledge or complex domain flows.
- 3 mismatched skills: Degraded performance (up to -10%) due to outdated assumptions or version incompatibilities conflicting with the repo.
Practical Implications for Harness Engineering
- Skills Are Code Interventions: Treat skills like code dependencies. Version them, evaluate them on paired benchmark tasks (with vs. without skill), and track token cost vs. pass-rate delta.
- Beware Prompt Bloat: Generic advice (e.g., “write clean modular code”) induces prompt bloat and increases inference latency without moving verification needles.
- Progressive Disclosure: Pre-load only skill names and triggers in the root prompt; load the full procedural body only when the task explicitly routes to that skill.
Related Concepts
- evals skills — Evaluating and testing agent skills.
- agent harness engineering — Harness architecture and runtime orchestration.
- prompt bloat — Degradation of performance from excessive prompt instructions.
- progressive disclosure — Staged context loading to avoid token saturation.
Evidence — verified primary sources
| raw/papers/skillsbench-2026.md | internal workspace doc | |
| raw/papers/swe-skills-bench-2026.md | internal workspace doc | |
| addy-osmani-agent-skills-2026 | https://addyosmani.com/blog/agent-skills/ | ingested 2026-08-27 sha256:8b3508b787f0… |
Graph context
References (4)
Evals Skills, SKILL.md) into an agent's context during inference. Empirical benchmarks demonstrate that skills are not a universal productivity booster;Prompt Bloatand increases inference latency without moving verification needles.Agent Harness EngineeringHarness architecture and runtime orchestration.Progressive DisclosureStaged context loading to avoid token saturation. Referenced by (1)
Agentic Code Qualityempirical evaluation of procedural skills