Measuring and graduating
This is the part that makes the rest defensible. Every rule's stance is a claim about your own behaviour, and the claim is checked against your own transcripts rather than asserted.
The loop is: mine → backtest → explain → review → compliance → graduate.
0. Mine — what is going wrong that no rule names?
Every step below prices a rule somebody already thought of. mine is the
step before that: it groups the transcripts by command shape and ranks the
shapes that went wrong.
amont-agent mine --since 2026-08-15
amont-agent mine --min-support 10 --min-rate 0.4
amont-agent mine --format cases >> tests/corpus/<new-rule>.cases
A shape is the parsed command with its literals masked — the branch, the path, the line number and the commit message taken out, the program, the verbs and the flag set kept:
git push origin feat/mine 2>&1 | tail -5 ─┐
git push origin fix/lint 2>&1 | tail -20 ─┴→ git push origin <word> 2>&1 | tail -5
A call counts as gone wrong when either of two things the transcript records is true:
- failed — the tool result carried
is_error: a non-zero exit, a refused permission, a run the harness killed; - corrected — a near-identical command followed within a few tool calls. A model that re-issues the same shape is a model whose first attempt did not land, and that is the interesting half: the mistakes worth a rule are the ones nothing reports.
Only the HEAD of a run of near-identical calls counts as a correction. A
polling loop is one decision repeated twenty times, not nineteen mistakes —
before that rule existed, the top line of the real report was echo waiting.
bad and rate are suspicion, not verdict. A shape re-run for good
reasons carries a high rate and names no mistake at all. That is why every
row prints its samples and why the next step is a person reading them.
Shapes an existing rule already fires on are listed apart, covered by <rule>, rather than proposed — and a covered shape that is still going wrong
is a rule that is observing when it should be advising.
What mine does not do is write the rule. Nothing learned here reaches
the hook: the path is mine → write the rule by hand → cases → corpus check → graduate, and what ships is the hand-written rule with its reason, its
remedy and its reviewed corpus. A guard that refused a command because a
clustering run found it suspicious could not explain itself to the person
whose work it just refused.
1. Backtest — what would this have cost me?
backtest replays your Claude Code transcripts through the rules and reports
firings per 1,000 tool calls per week, so a rule's cost is a number rather
than an impression.
amont-agent backtest --since 2026-07-06
amont-agent backtest --rule pipe-to-tail --json
amont-agent backtest --transcripts ~/.claude/projects # where they live
A weekly series is the thing to read, not a total. A habit that is halving on
its own does not need a deny; the model is already correcting. A flat line
over weeks is a habit that will not correct itself, and that is what promotion
is for.
2. Explain — look at the actual matches
A rate is only trustworthy if the matches behind it are real. explain prints
every match for one rule so you can read them.
amont-agent explain pipe-to-tail
amont-agent explain pipe-to-tail --sample 20
amont-agent explain pipe-to-tail --sample 20 --rank novelty
--rank novelty changes WHICH twenty. The default is the first twenty the
walk met — the oldest project, the oldest session, and, because a habit
repeats, very often twenty spellings of one command. Novelty picks the
twenty least like each other and least like the cases already in
tests/corpus/<rule>.cases, by greedy max-min over the shape distance, so
an hour of labelling buys as much of the precision estimate as an hour can.
It is deterministic: the same transcripts and the same corpus pick the same
cases, and a second pass does not hand back the first pass's.
3. Review — turn matches into reviewed judgements
Precision is kept as a corpus of judgements, not as a metric, because a metric charts a regression and a test prevents one.
amont-agent explain pipe-to-tail --format cases >> tests/corpus/pipe-to-tail.cases
$EDITOR tests/corpus/pipe-to-tail.cases # each `?` becomes match or nomatch
amont-agent corpus check # and this runs in the test suite
Include the cases that should not match. A corpus of positives alone measures recall and says nothing about how often the rule is wrong, which is the number that decides whether it can be allowed to refuse anything.
4. Compliance — is the advice worth its tokens?
advise buys its place in the model's context with tokens, every session,
forever. The backtest says what that costs. This says what it buys.
amont-agent backtest --compliance
amont-agent backtest --compliance --rule glob-in-flag-value --json
amont-agent backtest --compliance --window 40
For every firing it looks at what the model did next. If the next
equivalent command — same shape, within the window — no longer matches the
rule, the habit changed (complied). If it matches again, the advice was
read and ignored, or never reached the model (ignored). If nothing
equivalent followed, the firing is unanswered and counts towards neither.
Three things make the number readable:
- Per model. Models differ, and a pooled figure hides which one is
listening.
claude-opus-5andclaude-fable-5do not answer the same way to the same sentence. observerules are the control. They said nothing, so their share is the rate the habit corrects on its own. Advice is worth its tokens only where the advised share beats the silent one — the same argument the stance ladder rests on, measured instead of assumed.denyrules are left out. A refused command never ran, so there is no next command to compare it against.
What it cannot see: a transcript records what the model ran, not whether the hook spoke. A firing here means "the rule as it stands today would fire on this command", replayed. Where a rule has been widened since, or was observing then and advises now, the number is a reconstruction — and it is still the only evidence there is.
The Read tier from transcripts
The Read tool never reaches the backtester, so read-unbounded-large is
measured from the transcripts directly:
tools/read-rate.py
tools/read-rate.py --weeks 8
For each ISO week it prints the tool calls, the Reads over 16 KB with no
offset/limit after the same exemptions the rule makes (plans, diffs,
media, tool-results/), that count per 1000 calls, and the characters. Run
it to recount the rule's per_1000 and again two weeks after a release, to
compare against the weeks before.
The write tier from transcripts
lint-suppression-added never reaches the backtester either: an Edit or Write
is not a command. It is measured from the transcripts with
tools/suppression-rate.py:
tools/suppression-rate.py
tools/suppression-rate.py --weeks 8
tools/suppression-rate.py --until 2026-10-06
For each ISO week it prints the tool calls, the Edit, MultiEdit and Write calls
that add a suppression or loosen a lint setting, and that count per 1000 calls.
Run it to recount the rule's per_1000 and again two weeks after a release.
The script sees fragments only: a transcript holds what the model sent,
never the file, so it compares old_string with new_string, as the hook's
fragment fallback does. With no section around the text it undercounts the
config rows that need one ([tool.pyright], [tool.ruff*], [lints.*],
.golangci.yml disable: and exclude*), unless the edit carries the header
itself. --self-test classifies tests/fixtures/suppressions.txt, which a
Rust unit test classifies too; tests/suppression_rate.rs runs it.
What came before a publish, from transcripts
publish-without-skill asks a question the backtester cannot: whether the
tag-release or merge-when-green skill was called in this turn or the one
before. Its confirm reads the transcript, and the backtester never runs
confirm. It is measured with tools/skill-rate.py:
tools/skill-rate.py
tools/skill-rate.py --since 2026-10-07 --list
It prints the publishing Bash calls per 1000 (a v* tag pushed, gh pr merge, a merge sent to /pulls/<n>/merge), split by what preceded them:
no-skill, stale (called two or more prompts back), allowance (covered only
by the one prompt the window allows) and covered. It prints the same counts per
unique command per session, since a refused command is retried. Then it prints
how old the covering call was, and the uncovered rate by week. The matcher is a
regular expression. Against the Rust lexer it overcounts by about 6% (1,551
against 1,461 on 2026-10-07), so backtest --rule publish-without-skill remains
the count of matches. Run it again two weeks after a release.
5. Graduate — promote on the evidence
amont-agent graduate bare-stash-pop --to advise
amont-agent graduate bare-stash-pop --to deny
Promotion is gated on the corpus: a rule cannot be promoted past a corpus that does not support it.
Demotion is not gated at all
amont-agent demote bare-stash-pop
No questions, no evidence required, effective on the next command. This asymmetry is deliberate. A guard that is hard to back out of is one people uninstall instead of demoting — and uninstalling takes every rule with it, including the ones that were working.