Measuring and graduating

This is the part that makes the rest defensible. Every rule's stance is a claim about your own behaviour, and the claim is checked against your own transcripts rather than asserted.

The loop is: mine → backtest → explain → review → compliance → graduate.

0. Mine — what is going wrong that no rule names?

Every step below prices a rule somebody already thought of. mine is the step before that: it groups the transcripts by command shape and ranks the shapes that went wrong.

amont-agent mine --since 2026-08-15
amont-agent mine --min-support 10 --min-rate 0.4
amont-agent mine --format cases >> tests/corpus/<new-rule>.cases

A shape is the parsed command with its literals masked — the branch, the path, the line number and the commit message taken out, the program, the verbs and the flag set kept:

git push origin feat/mine 2>&1 | tail -5  ─┐
git push origin fix/lint  2>&1 | tail -20 ─┴→ git push origin <word> 2>&1 | tail -5

A call counts as gone wrong when either of two things the transcript records is true:

  • failed — the tool result carried is_error: a non-zero exit, a refused permission, a run the harness killed;
  • corrected — a near-identical command followed within a few tool calls. A model that re-issues the same shape is a model whose first attempt did not land, and that is the interesting half: the mistakes worth a rule are the ones nothing reports.

Only the HEAD of a run of near-identical calls counts as a correction. A polling loop is one decision repeated twenty times, not nineteen mistakes — before that rule existed, the top line of the real report was echo waiting.

bad and rate are suspicion, not verdict. A shape re-run for good reasons carries a high rate and names no mistake at all. That is why every row prints its samples and why the next step is a person reading them.

Shapes an existing rule already fires on are listed apart, covered by <rule>, rather than proposed — and a covered shape that is still going wrong is a rule that is observing when it should be advising.

What mine does not do is write the rule. Nothing learned here reaches the hook: the path is mine → write the rule by hand → cases → corpus check → graduate, and what ships is the hand-written rule with its reason, its remedy and its reviewed corpus. A guard that refused a command because a clustering run found it suspicious could not explain itself to the person whose work it just refused.

1. Backtest — what would this have cost me?

backtest replays your Claude Code transcripts through the rules and reports firings per 1,000 tool calls per week, so a rule's cost is a number rather than an impression.

amont-agent backtest --since 2026-07-06
amont-agent backtest --rule pipe-to-tail --json
amont-agent backtest --transcripts ~/.claude/projects   # where they live

A weekly series is the thing to read, not a total. A habit that is halving on its own does not need a deny; the model is already correcting. A flat line over weeks is a habit that will not correct itself, and that is what promotion is for.

2. Explain — look at the actual matches

A rate is only trustworthy if the matches behind it are real. explain prints every match for one rule so you can read them.

amont-agent explain pipe-to-tail
amont-agent explain pipe-to-tail --sample 20
amont-agent explain pipe-to-tail --sample 20 --rank novelty

--rank novelty changes WHICH twenty. The default is the first twenty the walk met — the oldest project, the oldest session, and, because a habit repeats, very often twenty spellings of one command. Novelty picks the twenty least like each other and least like the cases already in tests/corpus/<rule>.cases, by greedy max-min over the shape distance, so an hour of labelling buys as much of the precision estimate as an hour can. It is deterministic: the same transcripts and the same corpus pick the same cases, and a second pass does not hand back the first pass's.

3. Review — turn matches into reviewed judgements

Precision is kept as a corpus of judgements, not as a metric, because a metric charts a regression and a test prevents one.

amont-agent explain pipe-to-tail --format cases >> tests/corpus/pipe-to-tail.cases
$EDITOR tests/corpus/pipe-to-tail.cases    # each `?` becomes match or nomatch
amont-agent corpus check                   # and this runs in the test suite

Include the cases that should not match. A corpus of positives alone measures recall and says nothing about how often the rule is wrong, which is the number that decides whether it can be allowed to refuse anything.

4. Compliance — is the advice worth its tokens?

advise buys its place in the model's context with tokens, every session, forever. The backtest says what that costs. This says what it buys.

amont-agent backtest --compliance
amont-agent backtest --compliance --rule glob-in-flag-value --json
amont-agent backtest --compliance --window 40

For every firing it looks at what the model did next. If the next equivalent command — same shape, within the window — no longer matches the rule, the habit changed (complied). If it matches again, the advice was read and ignored, or never reached the model (ignored). If nothing equivalent followed, the firing is unanswered and counts towards neither.

Three things make the number readable:

  • Per model. Models differ, and a pooled figure hides which one is listening. claude-opus-5 and claude-fable-5 do not answer the same way to the same sentence.
  • observe rules are the control. They said nothing, so their share is the rate the habit corrects on its own. Advice is worth its tokens only where the advised share beats the silent one — the same argument the stance ladder rests on, measured instead of assumed.
  • deny rules are left out. A refused command never ran, so there is no next command to compare it against.

What it cannot see: a transcript records what the model ran, not whether the hook spoke. A firing here means "the rule as it stands today would fire on this command", replayed. Where a rule has been widened since, or was observing then and advises now, the number is a reconstruction — and it is still the only evidence there is.

The Read tier from transcripts

The Read tool never reaches the backtester, so read-unbounded-large is measured from the transcripts directly:

tools/read-rate.py
tools/read-rate.py --weeks 8

For each ISO week it prints the tool calls, the Reads over 16 KB with no offset/limit after the same exemptions the rule makes (plans, diffs, media, tool-results/), that count per 1000 calls, and the characters. Run it to recount the rule's per_1000 and again two weeks after a release, to compare against the weeks before.

The write tier from transcripts

lint-suppression-added never reaches the backtester either: an Edit or Write is not a command. It is measured from the transcripts with tools/suppression-rate.py:

tools/suppression-rate.py
tools/suppression-rate.py --weeks 8
tools/suppression-rate.py --until 2026-10-06

For each ISO week it prints the tool calls, the Edit, MultiEdit and Write calls that add a suppression or loosen a lint setting, and that count per 1000 calls. Run it to recount the rule's per_1000 and again two weeks after a release.

The script sees fragments only: a transcript holds what the model sent, never the file, so it compares old_string with new_string, as the hook's fragment fallback does. With no section around the text it undercounts the config rows that need one ([tool.pyright], [tool.ruff*], [lints.*], .golangci.yml disable: and exclude*), unless the edit carries the header itself. --self-test classifies tests/fixtures/suppressions.txt, which a Rust unit test classifies too; tests/suppression_rate.rs runs it.

What came before a publish, from transcripts

publish-without-skill asks a question the backtester cannot: whether the tag-release or merge-when-green skill was called in this turn or the one before. Its confirm reads the transcript, and the backtester never runs confirm. It is measured with tools/skill-rate.py:

tools/skill-rate.py
tools/skill-rate.py --since 2026-10-07 --list

It prints the publishing Bash calls per 1000 (a v* tag pushed, gh pr merge, a merge sent to /pulls/<n>/merge), split by what preceded them: no-skill, stale (called two or more prompts back), allowance (covered only by the one prompt the window allows) and covered. It prints the same counts per unique command per session, since a refused command is retried. Then it prints how old the covering call was, and the uncovered rate by week. The matcher is a regular expression. Against the Rust lexer it overcounts by about 6% (1,551 against 1,461 on 2026-10-07), so backtest --rule publish-without-skill remains the count of matches. Run it again two weeks after a release.

5. Graduate — promote on the evidence

amont-agent graduate bare-stash-pop --to advise
amont-agent graduate bare-stash-pop --to deny

Promotion is gated on the corpus: a rule cannot be promoted past a corpus that does not support it.

Demotion is not gated at all

amont-agent demote bare-stash-pop

No questions, no evidence required, effective on the next command. This asymmetry is deliberate. A guard that is hard to back out of is one people uninstall instead of demoting — and uninstalling takes every rule with it, including the ones that were working.