5K

Skill Activation as a Precision/Recall Measurement

Measures skill routing as classification with a labelled prompt dataset, per-skill precision and recall, and CI thresholds that fail the build

emerging Updated 2026-09-26 By Tomasz Religa (@uipreliga) Based on Coder Eval (UiPath) Complexity: medium Effort: days Impact: high
View .md Edit on GitHub
Add to Pack
or

Saved locally in this browser for now.

Cite This Pattern
APA
Tomasz Religa (@uipreliga) (2026). Skill Activation as a Precision/Recall Measurement. In *Awesome Agentic Patterns*. Retrieved September 28, 2026, from https://agentic-patterns.com/patterns/skill-activation-as-precision-recall
BibTeX
@misc{agentic_patterns_skill-activation-as-precision-recall,
  title = {Skill Activation as a Precision/Recall Measurement},
  author = {Tomasz Religa (@uipreliga)},
  year = {2026},
  howpublished = {\url{https://agentic-patterns.com/patterns/skill-activation-as-precision-recall}},
  note = {Awesome Agentic Patterns}
}

Use when

  • A catalog has more than a handful of skills, or more than one could plausibly answer a prompt
  • Skill descriptions are edited, merged, or added without a behavioural test

Avoid when

  • A single skill with no siblings to compete with
  • The agent's skill is invoked explicitly by name rather than selected by the model

Problem

A skill is usually tested as if it were a function: run it, check the output. But before a skill runs, the model must choose it. It reads your skill's description alongside every other skill's description and picks one. That choice is a routing decision, and it is what actually fails in production.

It fails in two directions, and both are silent:

  • Under-triggering (a recall problem). The skill never loads. Users experience this as "your feature does not work," and blame the feature.
  • Over-triggering (a precision problem). The skill loads on prompts it does not own. It hijacks a conversation, injects its playbook into an unrelated context, burns tokens on reference files nobody needed, and produces confidently wrong output because the agent is now steering toward the wrong goal.

The two are not symmetric in detectability. Under-triggering gets reported by users; over-triggering gets absorbed, because the output still looks like an answer. So the more damaging failure is the one you are less likely to hear about.

Worse, the decision is coupled across the whole catalog. Adding a skill, editing one sentence of a description, or merging two skills re-routes prompts that belong to skills you did not touch. Hand-testing five prompts once and shipping cannot see this, and nothing in a typical CI pipeline does either — descriptions are prose, so no linter, typechecker, or unit test has an opinion about them.

Solution

Treat activation as what it is — a classification problem — and measure it with the standard instrument: a labelled dataset, per-skill precision/recall/F1, a confusion view, and thresholds that fail the build.

1. Label a dataset, not a checklist. One row per prompt, tagged with the set of skills that should engage. Use an empty set when none should engage; include all legitimate skills for a multi-skill task:

{"id": "flow-001", "prompt": "Convert checkout to a .flow process", "expected_skills": ["acme-flow"]}
{"id": "ins-014", "prompt": "Why is my Insights dashboard empty?",  "expected_skills": ["acme-insights"]}
{"id": "neg-007", "prompt": "Set up Acme Insights and connect it to our warehouse", "expected_skills": []}

2. Include the negative region — it is most of real traffic. A positive-only dataset can still catch cross-skill interference (skill B firing on a row labelled for skill A is a real false positive). What it cannot see is the much larger space where no skill should fire. Those rows carry an empty expected set.

3. Evaluate every row against every skill. This is the step that makes the measurement whole. Do not ask "did skill A fire on its own rows?" — ask, for each of the N skills, what happened on all M rows. One agent trace per row, N verdicts from it:

for row in dataset:                      # M prompts
    trace = run_agent(row.prompt)        # one trace, reused N times
    for skill in catalog:                # N skills
        fired    = skill_engaged(trace, skill)     # tool call, or skill-file read
        expected = (skill in row.expected_skills)
        confusion[skill][expected][fired] += 1

4. Detect engagement harness-agnostically. Whichever way your agent records it — a native skill-invocation event, or the agent reading the skill's files off disk — normalize both to one boolean so the same dataset scores across harnesses.

5. Gate CI on the aggregate, not on individual rows. Individual rows are non-deterministic; the per-skill aggregate over a few hundred rows is stable enough to threshold:

success_criteria:
  - type: skill_triggered
    skill: acme-flow
    suite_thresholds: { recall: 0.70, precision: 0.90 }

Precision thresholds should usually be stricter than recall thresholds, because a false positive damages every interaction that brushes the trigger surface while a false negative degrades one skill.

graph TD A[Labelled prompts<br/>positive + negative rows] --> B[One agent run per prompt] B --> C[Which skills engaged?] C --> D[Per-skill confusion matrix] D --> E[precision / recall / F1] E --> F{Thresholds met?} F -- no --> G[Fail the build] F -- yes --> H[Ship]

Evidence

  • Evidence Grade: medium — one large internal suite plus several public catalogs measured; the technique is young and adoption is early.
  • Most Valuable Findings:
    • Routine edits move activation a lot. On a catalog of ~20 skills with ~1,000 labelled prompts, adding one sentence to each skill description moved micro-recall from 46.3% to 67.3%, and took the number of skills clearing 70% recall from 4 to 11. A change that large under a routine edit cannot be verified by hand once and then trusted.
    • Broad "router" descriptions absorb their own specialists. In a 19-row pilot against a public vendor catalog, a router skill whose description claimed the capabilities of four specialists fired on its own 2 prompts and on 3 belonging to specialists — precision 0.40. Every specialist that did fire, fired correctly (precision 1.00); every failure was recall. In 2 of the 3 absorbed rows the agent loaded the router and then reached for a web fetch instead of the local specialist — it absorbed the prompt and failed to route.
    • Collisions are findable statically, before any run. In the catalog above, the router's description claimed a superset of four specialists' own trigger lists, and one trigger phrase appeared verbatim in two different skills' descriptions — both readable off the frontmatter with no model calls. A static overlap scan is the cheap first pass, and it predicted the routing failures the run then measured.
  • Unverified / Unclear: how few rows suffice (≈50 shows signal; the stability floor is unmeasured); how much of the measurement transfers across model versions; whether precision or recall thresholds should dominate in a safety-critical catalog, where a missed diagnostic skill may outweigh a spurious one.

How to use it

Start here, in this order:

  1. Static scan first (free). Diff every pair of skill descriptions for shared trigger phrases and shared nouns. Overlap here predicts collisions and costs nothing to find.
  2. Fifty rows, one skill pair. Take the two skills you most suspect of competing. ~25 prompts each plus ~10 negatives will already show you a confusion matrix worth reading.
  3. Add the negative set. Sample prompts from real traffic that no skill should own. This is where over-triggering becomes visible.
  4. Widen to the catalog, then threshold. Once per-skill numbers are stable run to run, set thresholds slightly below current values and put them in CI, so the next description edit has to clear the bar it inherited.

Prerequisites: an agent runner that records which skills engaged; a sandbox, so rows cannot contaminate each other; and labels you trust. Labelling is the real cost — budget more time for deciding what each prompt should route to than for running the suite.

Practical notes:

  • State your labelling convention for out-of-scope prompts, and score them as their own group. A domain-agnostic question answered with zero tool calls is defensible behaviour, not a failure, and it will otherwise pollute your precision figure.
  • Give the agent enough turns. A too-tight turn budget reads as a non-activation.
  • Give codebase-shaped prompts a codebase. A prompt like "review the query patterns here" against an empty sandbox measures your fixture, not your routing.
  • Replicate. Two or three replicates per row separate a genuine routing failure from ordinary sampling noise.

Trade-offs

Pros:

  • Makes the silent failure mode visible, and the invisible one (over-triggering) measurable at all.
  • Catches catalog-wide coupling: the suite scores every skill on every row, so an edit to one description cannot quietly re-route another skill's prompts.
  • Produces a threshold that can gate CI, unlike a per-row assertion on a non-deterministic decision.
  • The same labelled dataset scores across harnesses, so it doubles as a harness comparison.
  • Directs the fix at the right artifact: descriptions, which are cheap to edit.

Cons:

  • Labelling is real work, and disagreements about what a prompt "should" route to are common and legitimate.
  • Costs model calls. One run per row, times replicates — the pilot above ran ~$0.06/row on a small model. Cheap per row, not free per catalog.
  • Non-deterministic, so only the aggregate is trustworthy; a single flipped row means nothing.
  • Version-bound. Numbers move with the model, so a threshold is a regression guard, not an absolute quality score.
  • Quadratic-ish reporting. N skills × M rows produces a lot of cells; it needs a confusion view to be readable.
  • Measures selection, not usefulness. A skill can route perfectly and still give bad advice. This pattern is upstream of output quality, not a substitute for it.

References