amont-agent
A guard that reads a shell command before Claude Code runs it, and can observe it, advise against it, or refuse it.
It exists because of a class of mistake no git hook can see, since the mistake is in the command string and never reaches a hook at all:
git push origin main 2>&1 | tail -5
A pipeline's exit status is the status of its last command. tail succeeds
at tailing an error message, so a rejected push, a push killed by a timeout,
and a push that never left the machine all report success — and the trimming
discards the error text too, so the failure is silent in both channels.
amont-agent is a PreToolUse hook. It is a single binary, it is opt-in, and
it phones nothing home.
The shape of it
Three stances, and the middle one is the point:
| stance | effect |
|---|---|
observe | records the firing and says nothing at all |
advise | puts the reason into the model's context; refuses nothing |
deny | refuses the tool call, with the reason and the remedy |
Of the 38 rules, three ship as deny, twenty-two as advise and thirteen as
observe. A rule is promoted only once your own transcripts say it should be
— see measuring and graduating — and amont-agent rules shows
the stance in force on your machine.
Start here
amont-agent install # prints the settings block, writes nothing
amont-agent install --write # merges it into ~/.claude/settings.json
amont-agent doctor # is it installed, runnable, and actually firing?
amont-agent status # every rule, its stance, and what it has seen
Relationship to amont
amont is a git-hook manager by the
same author. The two are independent: neither needs the other, they share no
code, and cargo tree here shows serde and serde_json and nothing else.
They meet in exactly one place, and it is optional. If amont is installed and
this repository carries its generated AGENTS.md block, a session opening on
a stale block is told so — see the session notice. With
no amont on PATH, that check says nothing.
Relationship to attest
attest is the CI end of the same story: amont signs a note at pre-push naming the gates that really ran, and attest verifies it so CI can skip them. It has no connection to this guard beyond the author and the conviction all three share — trust what was verified, never what was reported. The masked push at the top of this page is what that looks like when it fails.
Installing
The binary first, then the wiring. They are separate acts on purpose: a
program that can refuse the commands your agent runs should not install itself
into your settings.json as a side effect of you fetching it.
The binary
curl -fsSL https://raw.githubusercontent.com/fredericrous/amont-agent/main/install/install.sh | sh
irm https://raw.githubusercontent.com/fredericrous/amont-agent/main/install/install.ps1 | iex
Either one downloads a release binary, verifies it against the published
SHA256SUMS, and puts it in ~/.local/bin. Neither wires anything in.
Or: brew install fredericrous/tap/amont-agent, cargo install amont-agent,
or a binary straight from
Releases — Linux
x86_64 (gnu and static musl) and aarch64 (gnu), macOS (Intel and Apple
silicon), and Windows x86_64.
Why there is no npm package
There was one, briefly, and removing it is the more useful answer than listing install methods.
install bakes the ABSOLUTE path of the running binary into
settings.json, deliberately: a PATH-resolved command exits 127 into
Claude Code's non-blocking bucket the moment PATH differs, which disables
the guard with nothing to notice it. npm cannot supply a stable absolute
path. Under npx the binary lives in npm's _npx cache, which npm garbage-
collects; under a project-local npm i -D it lives in one project's
node_modules, while the guard it configures is machine-global — so
rm -rf node_modules in one repository would silently disable the guard for
every session on the machine.
amont has an npm package for a reason that does not transfer: it is a
per-repository tool, and npm i -D amont plus a prepare script means the
hooks travel with the repository. This is per-developer machine
configuration — ~/.claude/settings.json, global git config, a journal in
~/.claude/. Nothing about it belongs to a project.
Every channel above hands install a path that stays put.
The wiring
amont-agent install # prints the settings block, writes nothing
amont-agent install --write # merges it into ~/.claude/settings.json
install refuses to guess. It will not patch a settings.json it could not
parse, and it will not write one whose formatting it cannot reproduce — a diff
full of reformatting hides the one line it added — so it prints the block and
changes nothing unless --reformat says otherwise. uninstall removes exactly
what it wrote and leaves everything else byte-identical.
amont-agent install --write --project # .claude/settings.json instead
amont-agent install --write --local # .claude/settings.local.json
Four kinds of entry are written: the guard on PreToolUse, for Bash and for
the file tools; the assertions on PostToolUse for Bash, which check what a
command that reported success actually did; a PostToolUse entry for Read,
which remembers a file only once it has actually been read; and a
SessionStart entry that leaves a heartbeat and states where the checkout
stands against the remote. Without the heartbeat, doctor cannot tell "nothing
fired this week" from "the guard has been dead since Tuesday".
Your settings.json keeps the mode it already had, and a file created here
starts at 0600: it can hold MCP environment blocks, and those hold
credentials.
Knowing it is alive
Claude Code hooks fail open quietly: a command that cannot be resolved exits
127, which is a non-blocking status, and nothing tells you. doctor exits
non-zero when the guard is inert, so it can run from cron:
✓ installed in /Users/you/.claude/settings.json
✓ amont-agent 2.0.0 at /Users/you/.local/bin/amont-agent
✓ a refused command produces a valid decision document
✓ last ran 4m ago
✓ acting on pipe-to-tail
Turning it off
AMONT_AGENT_OFF=1 # this shell only
git config --global amont.agent.enabled false # everywhere
amont-agent uninstall --write # remove the settings entries
Demoting one rule is almost always the better move than switching the guard off — see stances. A guard that is hard to back out of is one people uninstall instead of demoting, and uninstalling takes every rule with it.
Stances
A rule does one of three things, and the middle one is the point:
| stance | effect |
|---|---|
observe | records the firing and says nothing at all |
advise | puts the reason into the model's context; refuses nothing |
deny | refuses the tool call, with the reason and the remedy |
observe and advise are not two ways of saying "not blocking yet".
additionalContext enters the model's context and therefore changes its
behaviour, which contaminates the rate the observation exists to measure. A
rule that talks is intervening. That is why the two are named differently
and why backtest numbers from an advise rule are not comparable with the
numbers that justified promoting it.
That contamination is also what makes it measurable. amont-agent backtest --compliance reads the two stances against each other: what a model did
after an advise fired, beside what it did after an observe fired and
said nothing. The second is the rate the habit corrects on its own, and an
advise that does not beat it is costing context tokens for nothing. See
measuring and graduating.
Changing one
Takes effect on the next command; nothing to restart.
git config --global amont.agent.pipe-to-tail.stance observe
git config --global amont.agent.stance observe # every rule
The ladder, most specific first:
rule.default_stance < amont.agent.stance < amont.agent.<id>.stance
then capped at the rule's own ceiling, and clamped to observe if the guard
is switched off.
Ceilings
Every rule declares the loudest stance it may ever take. For most rules that
is deny, and the ceiling changes nothing. A rule whose finding is an
estimate — "this loop may make hundreds of requests" — is capped at
advise: refusing a command on an estimate would claim a certainty the
analysis does not have.
The cap is applied after every configured key, so neither
amont.agent.stance deny nor the rule's own key can pass it, and
graduate --to deny refuses a capped rule with the reason. amont-agent rules prints each rule's ceiling beside its stance.
Why git config and not a committed file
Promotion power stays on the machine, with the person. A rule that a committed
file could promote to deny would mean cloning a repository hands it the
power to refuse your shell commands.
This is enforced three times, because neither of the first two was enough.
amont.conf's parser will not name these rules, so a committed manifest cannot reach them.- amont's config reader lets a repository's committed policy
setlines outrank system and global git config — right for a hook manager, wrong for a guard — so this project reads git config itself, without that ladder. - Git's own search order ends at
--localand--worktree, and the last file wins. A.git/configtherefore outranked your--globalanswer without any of amont's machinery being involved — and the agent whose command was just refused can write one, sincegit configis not a command any rule here objects to. So the reader takes--global, then--system, and nothing else.
A stance answers to your own git config and to nothing a repository carries or
a process standing in one can write. A file your global config includes or
includeIfs counts as your own: git skips those for a scoped read unless asked,
and this reader asks. graduate and demote write --global
for the same reason.
One consequence worth stating plainly: git config amont.agent.<rule>.stance
run inside a repository writes --local by default, and this tool will not
read it. Pass --global, which is what every example here does.
The rules
| rule | ships as | what it catches |
|---|---|---|
pipe-to-tail | deny | a mutating command whose status is swallowed by a pipe |
bare-stash-pop | observe | git stash pop with no ref, where refs/stash is shared across worktrees |
gh-pr-merge-auto | observe | --auto on a repository with no required checks, which merges immediately |
publish-without-skill | advise | a v* tag pushed or a pull request merged with no call of the tag-release / merge-when-green skill in this turn or the one before, so the procedure ran from memory |
forge-merge-by-hand | advise | merging a pull request by POSTing to the forge's merge endpoint, which answers 200 whether the checks passed, failed or never started |
forge-status-stale-row | observe | keeping the first row of a commit's append-only /statuses list, so a stale pending reads as the present and the wait never ends |
no-verify | observe | turning the whole commit gate off rather than one check |
git-add-broad | observe | staging the tree instead of the change |
stale-base | advise | a branch or worktree started from a checkout the remote has moved past |
push-preflight | advise | a git push whose slow pre-push test gate has not been rehearsed with amont rehearse --wait |
push-preview | advise | a push that would publish interface changes no approved localhost preview covers, unless the plan the person approved declares the evidence is enough (preview approval) |
plan-review-panel | deny | a plan presented at ExitPlanMode before its expert review panel ran (the review panel) |
plan-phases-open | deny | a turn ending while the plan this branch carries still has an open phase that is not a 🧑 decision: — the agent is sent on to it |
implementation-review | advise | a push of a branch that carries a plan, whose diff no independent reviewer has read for the tree being pushed (the implementation review) |
foreground-poll | advise | a polling loop or gh run watch in the foreground, where the tool's ten-minute clock will kill it one poll short |
unbounded-background-push | advise | a git push sent to the background with its output kept in a file and no timeout, so the pre-push gate inside it can run for as long as it likes with nothing ending it or reporting it |
sed-in-place | advise | sed -i spelled for the other sed (-i '' on GNU, bare -i on BSD) |
kubectl-gitops | advise | an imperative kubectl write in a repository Flux or Argo reconciles |
tag-after-commit | advise | git tag chained onto a git commit that a hook may have refused |
release-tag-push | observe | pushing a v* tag, which publishes: the tag may name the wrong commit, and a green workflow does not mean the artefact is right |
worktree-remove-force | advise | git worktree remove --force on a worktree that still holds uncommitted work |
amend-pushed | advise | git commit --amend on a commit the remote already has |
branch-force-delete | observe | git branch -D on a branch whose commits are on no remote and not merged |
poll-blank-verdict | observe | a wait that stops on any value but the one it names, so a failed lookup's empty string reads as the answer |
worktree-isolation | observe | a branch created, or git reset --hard, in the primary checkout of a repository that already has linked worktrees |
stdin-hang | observe | a command that will read standard input with nothing on it — cat > file, a bare interpreter, tee outside a pipe — which blocks silently until the tool's clock runs out |
glob-in-flag-value | advise | an unquoted glob inside a flag value (--include=*.ts), which zsh expands — or fails on — before the program sees it |
glob-no-match | advise | an unquoted glob operand that matches nothing, which under zsh aborts the clause before it starts while a later clause reports success |
equals-separator | observe | a bare word beginning with = (echo ===, [ x == y ]), which zsh reads as a command lookup and whose failure aborts the whole command list |
unsplit-expansion | observe | an unquoted $var under set -- or for … in, which zsh leaves as one word: the positionals or the loop get the whole string and the command runs wrong while reporting success |
path-operand-missing | advise | a read of a path that is not there — under the grep shim a warning in mid-stream, for the coreutils one line and carry on — which the chain then reports as success |
stat-bsd-format | advise | stat -f '%…' on GNU stat (or -c on BSD), which prints a filesystem report where a timestamp was wanted |
whole-file-dump | advise | a file poured whole into the tool result by cat/sed -n/head — 31% of all result bytes measured — where the Read tool would have windowed it |
persisted-output-dump | advise | reading back whole a tool result the harness saved to a file for being too large, paying for it twice |
file-reread | advise | a Read, or a cat, of a file this session already has in context and that is unchanged on disk since — answered from the session's own record, not the command |
read-unbounded-large | advise (ceiling deny) | a Read with no offset/limit of a file over 16 KB, where the whole file lands in the context and every later turn carries it; plan files and diffs a reviewer is handed are exempt |
lint-suppression-added | advise (ceiling deny) | an Edit, MultiEdit or Write that adds a lint suppression (# type: ignore, # noqa, eslint-disable, @ts-ignore, #[allow], //nolint) or loosens a lint configuration (tsconfig, eslint, pyright, ruff, Cargo [lints], golangci), where general.no-disabled-safety says to fix the finding instead; the file is compared before and after by marker, not by line |
request-fanout | observe (ceiling advise) | one command that may make more than 50 explicit network transfers to one destination — a loop following Link: next, gh api --paginate, a curl URL range — counted by the shell analysis (analysis.md) with the loops and calls that multiply them |
Two more checks run after a command rather than before it, on PostToolUse:
push-landed and push-published verify what a push that reported success
actually did. See the assertions.
amont-agent rules prints this with each rule's measured firing rate, and
with the stance in force on this machine rather than the shipped one: a rule
you promoted in ~/.gitconfig reads deny (ships as observe), and a rule
with a ceiling names it, (max advise).
Why only three of them deny
pipe-to-tail blocks because seven consecutive weeks of measurement showed no
downward trend while every other habit halved. That is the bar: a rule earns
deny from your own transcripts, not from an argument about how bad the
mistake is. A habit the model is already correcting does not need a deny.
See measuring and graduating.
plan-review-panel is the other kind of deny, and the person's choice
rather than a measurement: it fires on ExitPlanMode, not on a shell command,
and checks a fact — whether the review panel the plan calls for ran on the plan
being presented — naming the missing roles when it did not. After two refusals
of the same plan, or whenever it cannot check, it hands the call to the person
as ask instead. See the review panel.
plan-phases-open is the third, also the person's choice. It fires on
Stop, the moment the agent ends its turn. ADR-0022 makes a plan one branch
and one pull request with its phases as commits, so a turn that ends with
a phase still open is a stall, not a handoff. The rule reads only the
active plans changed on the current branch (against the merge-base with
the remote default branch), never one already on main, and sends the agent
on to the first open - [ ] under ## Phases. It is silent when that
phase is a 🧑 decision:, in plan mode, while background tasks are still
running, and when the agent's last message has a line starting with
WAITING: <reason> (a preview awaiting approval, a deferred phase, a
blocker). After three continuations on the same phase in one session it lets
the turn end and tells the person, with a note only they see; their next
prompt resets the count. Anything it cannot establish — no repository, a git
failure, an unreadable plan, a session id it will not use as a file name —
lets the turn end. To turn it off without uninstalling:
git config --global amont.agent.plan-phases-open.stance observe.
stale-base advises from the start because it refuses nothing, speaks only
after measuring a real gap, and names a failure no correcting loop can see —
nothing fails when you build on stale code. The work is correct against
the code it can see, and the conflict arrives later, from somewhere else. So a
session opening in a checkout the remote has moved past is told — see
the session notice for why it fetches and never pulls.
push-preflight advises for the same reason. git opens its connection to the
remote before it runs pre-push and holds it idle for as long as the test
gate takes; a remote that closes idle sessions kills the push after the gate
has already passed, and the model reads "the network" where the cause was the
gate's placement. With amont ≥ 1.28, amont rehearse --wait runs the same
gate on a snapshot of HEAD with no connection open — or follows the
rehearsal amont.rehearseOnCommit already started — and stamps the tree, so
the push that follows skips the suite (amont run pre-push on 1.27). The rule
speaks only when confirm finds all three facts: amont guards this
repository's pushes, a test gate would run for this push, and HEAD's tree
carries no stamp yet.
Shape, then the world
Most rules after the first handful came out of the transcripts the same way —
tens of thousands of Bash calls, sorted by what failed, was killed, or drew a
correction (see mine).
Each fires on shape and, where the fact lives in the world, confirms it first:
whether the call already runs in the background, which sed is on PATH,
whether the repository holds a Flux or Argo resource, whether the worktree is
dirty, whether the remote has the commit, whether any other branch has the
commits. A rule that fires on shape alone, like tag-after-commit, names a
failure every command in the chain reports as success.
pipe-to-tail in full
git push origin main 2>&1 | tail -5
A pipeline's exit status is its last command's. tail succeeds at tailing
an error message, so a rejected push, a push killed by a timeout, and a push
that never left the machine all report success — and the trimming discards the
error text, so the failure is silent in both channels.
The remedy the rule prints is not "don't use tail": it is to run the mutating
command on its own, read its output afterwards, and then verify the effect
(git ls-remote origin refs/heads/<branch>) rather than the exit code.
set -o pipefail first, or writing to a file and tailing the file, are also
accepted — the rule fires on the shape, and the reason names all three ways
out.
Asking about one command
amont-agent check 'git push | tail -1'
No stdin, no session, no journal entry — just the rules over one string, with whatever they would have said.
The assertions
A rule reads a command before it runs. An assertion checks what a command that already reported success actually did.
The two halves are not the same job. pipe-to-tail can refuse
git push … | tail -5 because the mistake is visible in the command string.
Nothing in a command string tells you that the push you just ran reported
Everything up-to-date about a branch you were not on.
| id | fires on | asks |
|---|---|---|
push-landed | git push | is the branch on the remote, at the commit you have? |
push-published | git push to a UI repository | did it publish a new commit, and was it an approved preview? (records only; see preview approval) |
Only successful calls
Claude Code sends a failed tool call to PostToolUseFailure, an event this
crate ignores. So everything an assertion sees claimed to work — which is the
whole point. A command that returned an error is already in front of the model;
there is nothing invisible left to point out.
An assertion cannot refuse
The tool has already run. A deny stance therefore speaks exactly like
advise, the same way it does at a session opening.
What an assertion can do is state a fact:
amont-agent/push-landed: `git push` exited 0, but origin/main is at d0765d4f9
while the local branch is at 8970c7ffb. The push did not land. Run the push
again on its own and read its output, then confirm with `git ls-remote`.
That is not advice to weigh. It is the remote's answer.
What it refuses to judge
push-landed handles the unambiguous shapes and nothing else. A
HEAD:refs/heads/other refspec, a tag push, several refspecs at once,
--delete, --mirror, --all: each needs a different question asked of the
remote, and a confidently wrong accusation costs the channel its credibility —
which is the only thing the channel has.
The same goes for everything it cannot establish: a detached HEAD, a remote given as a URL, a working directory that has gone, a remote that cannot be reached without a password. All silence.
It can never prompt
Hooks run with no controlling terminal, so a credential prompt does not fail —
it hangs, and it hangs the session rather than this process. Every child runs
with GIT_TERMINAL_PROMPT=0, GIT_ASKPASS/SSH_ASKPASS disabled, ssh in
BatchMode, and a five-second deadline. A guard installed to make pushing safer
must never be the reason a push becomes impossible.
Measuring
examine is pure, so the backtester replays it like any rule:
$ amont-agent backtest push-landed --since 2026-07-06
push-landed 1007 38.7 routine observe
Read that number correctly. For a rule, a firing is a mistake caught. For an
assertion, a firing is a question asked — 1,007 pushes in 29,758 Bash calls,
about one call in twenty-six paying for one git ls-remote. It speaks only when
the answer disagrees, which is far rarer.
How often a claim is actually broken is not backtestable at all: verify
touches the world, and the world has moved since those commands ran. That number
accumulates forward, in the journal:
$ grep push-landed ~/.claude/amont-agent/journal.log
held is a push that landed, broken one that did not, unverified one this
crate declined to judge.
Measuring and graduating
This is the part that makes the rest defensible. Every rule's stance is a claim about your own behaviour, and the claim is checked against your own transcripts rather than asserted.
The loop is: mine → backtest → explain → review → compliance → graduate.
0. Mine — what is going wrong that no rule names?
Every step below prices a rule somebody already thought of. mine is the
step before that: it groups the transcripts by command shape and ranks the
shapes that went wrong.
amont-agent mine --since 2026-08-15
amont-agent mine --min-support 10 --min-rate 0.4
amont-agent mine --format cases >> tests/corpus/<new-rule>.cases
A shape is the parsed command with its literals masked — the branch, the path, the line number and the commit message taken out, the program, the verbs and the flag set kept:
git push origin feat/mine 2>&1 | tail -5 ─┐
git push origin fix/lint 2>&1 | tail -20 ─┴→ git push origin <word> 2>&1 | tail -5
A call counts as gone wrong when either of two things the transcript records is true:
- failed — the tool result carried
is_error: a non-zero exit, a refused permission, a run the harness killed; - corrected — a near-identical command followed within a few tool calls. A model that re-issues the same shape is a model whose first attempt did not land, and that is the interesting half: the mistakes worth a rule are the ones nothing reports.
Only the HEAD of a run of near-identical calls counts as a correction. A
polling loop is one decision repeated twenty times, not nineteen mistakes —
before that rule existed, the top line of the real report was echo waiting.
bad and rate are suspicion, not verdict. A shape re-run for good
reasons carries a high rate and names no mistake at all. That is why every
row prints its samples and why the next step is a person reading them.
Shapes an existing rule already fires on are listed apart, covered by <rule>, rather than proposed — and a covered shape that is still going wrong
is a rule that is observing when it should be advising.
What mine does not do is write the rule. Nothing learned here reaches
the hook: the path is mine → write the rule by hand → cases → corpus check → graduate, and what ships is the hand-written rule with its reason, its
remedy and its reviewed corpus. A guard that refused a command because a
clustering run found it suspicious could not explain itself to the person
whose work it just refused.
1. Backtest — what would this have cost me?
backtest replays your Claude Code transcripts through the rules and reports
firings per 1,000 tool calls per week, so a rule's cost is a number rather
than an impression.
amont-agent backtest --since 2026-07-06
amont-agent backtest --rule pipe-to-tail --json
amont-agent backtest --transcripts ~/.claude/projects # where they live
A weekly series is the thing to read, not a total. A habit that is halving on
its own does not need a deny; the model is already correcting. A flat line
over weeks is a habit that will not correct itself, and that is what promotion
is for.
2. Explain — look at the actual matches
A rate is only trustworthy if the matches behind it are real. explain prints
every match for one rule so you can read them.
amont-agent explain pipe-to-tail
amont-agent explain pipe-to-tail --sample 20
amont-agent explain pipe-to-tail --sample 20 --rank novelty
--rank novelty changes WHICH twenty. The default is the first twenty the
walk met — the oldest project, the oldest session, and, because a habit
repeats, very often twenty spellings of one command. Novelty picks the
twenty least like each other and least like the cases already in
tests/corpus/<rule>.cases, by greedy max-min over the shape distance, so
an hour of labelling buys as much of the precision estimate as an hour can.
It is deterministic: the same transcripts and the same corpus pick the same
cases, and a second pass does not hand back the first pass's.
3. Review — turn matches into reviewed judgements
Precision is kept as a corpus of judgements, not as a metric, because a metric charts a regression and a test prevents one.
amont-agent explain pipe-to-tail --format cases >> tests/corpus/pipe-to-tail.cases
$EDITOR tests/corpus/pipe-to-tail.cases # each `?` becomes match or nomatch
amont-agent corpus check # and this runs in the test suite
Include the cases that should not match. A corpus of positives alone measures recall and says nothing about how often the rule is wrong, which is the number that decides whether it can be allowed to refuse anything.
4. Compliance — is the advice worth its tokens?
advise buys its place in the model's context with tokens, every session,
forever. The backtest says what that costs. This says what it buys.
amont-agent backtest --compliance
amont-agent backtest --compliance --rule glob-in-flag-value --json
amont-agent backtest --compliance --window 40
For every firing it looks at what the model did next. If the next
equivalent command — same shape, within the window — no longer matches the
rule, the habit changed (complied). If it matches again, the advice was
read and ignored, or never reached the model (ignored). If nothing
equivalent followed, the firing is unanswered and counts towards neither.
Three things make the number readable:
- Per model. Models differ, and a pooled figure hides which one is
listening.
claude-opus-5andclaude-fable-5do not answer the same way to the same sentence. observerules are the control. They said nothing, so their share is the rate the habit corrects on its own. Advice is worth its tokens only where the advised share beats the silent one — the same argument the stance ladder rests on, measured instead of assumed.denyrules are left out. A refused command never ran, so there is no next command to compare it against.
What it cannot see: a transcript records what the model ran, not whether the hook spoke. A firing here means "the rule as it stands today would fire on this command", replayed. Where a rule has been widened since, or was observing then and advises now, the number is a reconstruction — and it is still the only evidence there is.
The Read tier from transcripts
The Read tool never reaches the backtester, so read-unbounded-large is
measured from the transcripts directly:
tools/read-rate.py
tools/read-rate.py --weeks 8
For each ISO week it prints the tool calls, the Reads over 16 KB with no
offset/limit after the same exemptions the rule makes (plans, diffs,
media, tool-results/), that count per 1000 calls, and the characters. Run
it to recount the rule's per_1000 and again two weeks after a release, to
compare against the weeks before.
The write tier from transcripts
lint-suppression-added never reaches the backtester either: an Edit or Write
is not a command. It is measured from the transcripts with
tools/suppression-rate.py:
tools/suppression-rate.py
tools/suppression-rate.py --weeks 8
tools/suppression-rate.py --until 2026-10-06
For each ISO week it prints the tool calls, the Edit, MultiEdit and Write calls
that add a suppression or loosen a lint setting, and that count per 1000 calls.
Run it to recount the rule's per_1000 and again two weeks after a release.
The script sees fragments only: a transcript holds what the model sent,
never the file, so it compares old_string with new_string, as the hook's
fragment fallback does. With no section around the text it undercounts the
config rows that need one ([tool.pyright], [tool.ruff*], [lints.*],
.golangci.yml disable: and exclude*), unless the edit carries the header
itself. --self-test classifies tests/fixtures/suppressions.txt, which a
Rust unit test classifies too; tests/suppression_rate.rs runs it.
What came before a publish, from transcripts
publish-without-skill asks a question the backtester cannot: whether the
tag-release or merge-when-green skill was called in this turn or the one
before. Its confirm reads the transcript, and the backtester never runs
confirm. It is measured with tools/skill-rate.py:
tools/skill-rate.py
tools/skill-rate.py --since 2026-10-07 --list
It prints the publishing Bash calls per 1000 (a v* tag pushed, gh pr merge, a merge sent to /pulls/<n>/merge), split by what preceded them:
no-skill, stale (called two or more prompts back), allowance (covered only
by the one prompt the window allows) and covered. It prints the same counts per
unique command per session, since a refused command is retried. Then it prints
how old the covering call was, and the uncovered rate by week. The matcher is a
regular expression. Against the Rust lexer it overcounts by about 6% (1,551
against 1,461 on 2026-10-07), so backtest --rule publish-without-skill remains
the count of matches. Run it again two weeks after a release.
5. Graduate — promote on the evidence
amont-agent graduate bare-stash-pop --to advise
amont-agent graduate bare-stash-pop --to deny
Promotion is gated on the corpus: a rule cannot be promoted past a corpus that does not support it.
Demotion is not gated at all
amont-agent demote bare-stash-pop
No questions, no evidence required, effective on the next command. This asymmetry is deliberate. A guard that is hard to back out of is one people uninstall instead of demoting — and uninstalling takes every rule with it, including the ones that were working.
Preview approval
push-preview and push-published carry ADR-0028 (work.preview-approval,
fleet decision corpus, replacing ADR-0023): a commit that changes a user
interface is published only after the person approved that commit, seen
on localhost — unless the plan the person approved declares the evidence
is enough (below, rule
work.preview-unless-planned-evidence).
The rule exists because interface regressions kept reaching pull requests and deployed sites after the agent had checked its own screen. The person's eyes are the check that does not share the agent's blind spots, and localhost is where rejecting a screen costs one sentence instead of a fix pull request and a release.
The flow
The person works on several projects at once and does not remember where
this one stood. A bare "[preview f0223d3] Ship website-builder@f0223d3?"
gives them nothing to decide with, so a preview reaches them as a guide
(fleet rule work.preview-is-guided), three ways at once: the brief in the
terminal, a local page, and the app already open in their browser.
-
Verify. The agent drives the changed route in a real browser, checks the console and network, and takes before and after screenshots.
-
Write the guide — markdown, outside the worktree (
~/.claude/amont-agent/attestations/<sha>/guide.mdby convention, the screenshots beside it), because a file inside it, or the page rendered beside it, would dirty the commit it describes. The format is below. -
Serve and register. From the clean worktree, at the commit to be pushed:
amont-agent preview register --url http://localhost:5173/settings \ --guide ~/.claude/amont-agent/attestations/abc1234/guide.mdRun it in the foreground as the last command of its line. Anything may come before it, joined by
&∨:cd ~/Developer/app-wt-settings && npm run build && amont-agent preview register --url … --guide …It refuses a dirty tree, a non-commit
HEAD, a guide inside the worktree and a guide missing a required section (exit 1; stderr lists exactly what is missing). Otherwise it renders the guide toindex.htmlbeside it and prints one JSON object, compact on one line:field what it is idthe registration; it starts with the commit's sha7 repo,committhe checkout's top level and its full HEADurlwhere the clean worktree is served guide,page,page_urlthe guide, its rendered index.html(absolute), and that page as afile://URLlabel,aliasesthe names a person knows the commit by: the checkout's directory, then the main worktree's directory and the origin repository's name, each @sha7question_prefix[preview <id>] <label>: what the marked question starts with, verbatimattestationguideagain, for hooks older than 2.21.0--openalso hands the page to the platform opener (open,xdg-open,cmd /c start) without waiting; failing to open never fails the command.The
PostToolUsehook binds that output to the session and to the person's current prompt, in the repository the command ran in — the leadingcds followed — which must still be the printed repository at the printedHEAD; the journal line names the page. The JSON is read from the last non-empty line of stdout, so what earlier clauses print does not matter. The register must be the last clause: one that is followed by anything (register && echo), piped (| tee), redirected, inside$(…), after||, or in the background is not bound, and the journal says why.A line an earlier clause printed cannot pass for the register's own (
printf '<json>'; amont-agent preview register --bad): thePreToolUsehook stamps when the call started, inpreviews/registers/<tool_use_id>, and the bind requires that the id starts with the commit's sha7 and that the page exists, names the full commit and was written after that stamp. A call with no stamp is not bound. Stamps nobody came back for are swept after a day. -
Open the app and the page. Before asking, the agent opens the app in the person's browser at the state the first step reaches (the panel already open, the form already filled) and the rendered page beside it. That is the agent's job — the
worktree-taskskill's F7, with Claude in Chrome — not this binary's:registeronly renders the page and, with--open, opens it. -
Ask, in the same turn. The brief first — the guide's sections, in the terminal — then one
AskUserQuestionwhose text starts with the register'squestion_prefix([preview <id>] <label>), with options exactlyApprove,Request changes,Hold— no "(Recommended)" suffix. -
Push after
Approve. The approval covers that commit; a new commit, amend or rebase needs a new preview.
--attestation <file> is accepted for 2.21.0 only, as a deprecated alias of
--guide: the file is read as the guide and refused unless it is one.
The guide
Markdown with these five H2 sections. Headings match with emoji,
punctuation and case aside (## 📍 Where we are: is Where we are); the
page shows them in this order, then any other section the guide adds.
| section | what it carries | checked |
|---|---|---|
Where we are | project, branch and plan, and why the person is asked. Its first line is the page's title | present, not empty |
What you should see | in their words; "nothing new" when that is the point | present, not empty |
Try it | numbered steps: the exact URL, each click named by its visible label and position, and after each step what should appear | at least one numbered step (1.) and one http(s) URL |
Reference | before and after screenshots | present, not empty |
Already checked | what the agent verified, and what to look at especially | present, not empty |
A full example (guide.md, with before.png, after.png and panel.png
beside it):
# Settings: Save moves to the header
## 📍 Where we are
duro-app, branch `feat/settings-save`, plan **Settings panel** (phase 2 of 3).
You are asked because the Save button moved; nothing else on the page changed.
## What you should see
The **Save** button now sits in the header, top right, instead of at the
bottom of the form. Saving works exactly as before.
## 👉 Try it
1. Open http://localhost:5173/settings
- You should see the Settings form, with **Save** in the header, top right.
2. Change **Display name** (first field) to anything.
- **Save** turns from grey to blue.
3. Click **Save**, top right.
- A green toast "Saved" appears bottom left, and **Save** is grey again.

## Reference


## Already checked
- `/settings` at 1280 and 390 px wide: no console error, no failed request.
- Keyboard: Tab reaches **Save** right after the page title.
- Look especially at the narrow window: the button must not cover the title.
The page is one self-contained file: inline CSS readable in light and dark
(prefers-color-scheme), no script, no external asset. Its title is
repo@shortsha — <first line of Where we are>; under it, a large Open the
app link to --url, then the sections. Images (,
relative to the guide's directory) render inline; a missing one is a warning
on stderr, not a refusal. When one image's file name says before and
another's says after, the two show side by side (class="pair") at the top
of Reference. The converter is hand-rolled for this subset — headings,
paragraphs, bullet and numbered lists with one nested level, fenced code,
quotes, bold, code, links, images and bare URLs — and escapes
everything else; a link whose scheme is not http(s), mailto or file is text.
No markdown crate: amont-agent is on a trust path and a page only the person
reads does not clear change.dependency-bar.
Mockup mode
When the previewed branch commits a picked mockup — a *.dc.html artboard
under docs/mockups/<screen>/ (or the directory git config amont.agent.preview.mockups names; the application-landscape ui-handoff
PR check reads the same key) — the guide must prove fidelity
(handoff.prove-fidelity, ADR-0016). Added to the five sections:
- under
## Reference: the screen's path, aViewport: 1120px · lightline (the width and theme both images were taken at), and the picked artboard beside the built screen as image files next to the guide:artboard.pngandafter.png(mockupandliveare read the same; states pair by suffix,artboard-empty.pngwithafter-empty.png); - a section
## Differences from the mockup:None, or one bullet per difference, each- fixed: …or- deliberate: <reason>.
The range is the branch's own commits, from its merge base with the default
branch — never from its upstream, which after a first push already holds the
artboards. Artboards on a branch whose base cannot be resolved are a refusal,
not a skip. A later branch that only edits a screen whose artboards are
already on main is not in mockup mode; the PR check asks it for the
side-by-side or Mockup: none — <reason>.
The images are copies of the proof PNGs committed next to the artboards, so
copying them beside the guide comes first, as its own command: a
register chained after a cp is not bound (see below). Two PNGs whose widths
differ by more than a tenth are a warning. On the page, each pair shows full
width at the same scale, each image links to its full-size file, and an
Overlay toggle lays the built screen over the mockup at half opacity.
A branch in mockup mode that is pushed without a bound approval is held
(deny) rather than advised, unless the rule is set to observe.
What approves
| the person… | result |
|---|---|
picks Approve on the marked question | every pending registration whose id the marker lists is approved |
picks Request changes or Hold | the previews are dropped |
| does not answer (a timeout) | still pending |
types approve, ship, lgtm or looks good as the next prompt, after a marked question was asked | approved |
types anything else next — including yes, go, ok, continue | the previews lapse |
The id in the marker is what approves: it begins with the commit's sha7,
which the person sees, and the registration it names is already bound to
the session, the prompt and the question. The label is for the person; a
question without it still approves (it did not before 2.30.0, and the
person was asked twice). A marker that names no pending id approves
nothing. A register that did not bind — not the last command, run in the
background, with no PreToolUse stamp, or whose output does not match
HEAD or this call — says so right after it runs, with the command to run
and the question to ask.
Unmarked questions are ignored, so a release question answered "Approve" approves no preview, even in the same turn. Another session's answers never touch this session's previews. Registrations and approvals lapse after a day, checked whenever they are read.
AskUserQuestion accepts answers in its input, so a model could pre-answer
its own question. The PreToolUse hook records, per tool_use_id, whether a
marked question arrived pre-answered; such a question never marks its
previews as asked (so no answer to it can approve), the PostToolUse side
checks the same record again, and under deny the question is refused before
it runs.
When the rule speaks
push-preview fires on a push only when all of these hold:
- the repository has a user interface: a
devscript inpackage.jsonat the root or inweb/, orgit config amont.agent.push-preview.ui true(falseopts a repository out); - a pushed branch carries an interface change (below);
- the plan the branch landed does not declare the evidence (below);
- that commit has no approval.
Planned evidence
Some previews leave the person nothing to judge: they already decided the only visible change when they approved the plan (ADR-0028). The exemption is declared in the plan's body, so the person decides it when they approve the plan:
## Preview
evidence: the only visible change is the 28px small controls, decided in this plan.
Two hooks make it checkable:
-
The approval record. A
PostToolUsehook onExitPlanModereadstool_input.planFilePathwhen it runs (the person may have edited the plan) and, when the result has the approve shape (tool_responsean object carryingplan, nois_error), writes the plan's canonical body sha (plan review: front matter and review sections left out) to~/.claude/amont-agent/plan-approved/<sha>, journalledplan-approved. -
The check, at the push. For a branch whose interface change carries no picked mockup (mockup mode always asks; a screen list that cannot be taken is held as before), the evidence stands when all hold:
- the default branch is known (
checkout.defaultRemoteor the single remote, then itsHEAD,mainormaster) and shares a merge base with the pushed commit; - the branch adds a plan under
docs/plans/since that merge base (the first one added), and it is not a pointer (canonical:); - that plan, read at the commit that added it, has its canonical sha
in
plan-approved/; - its
## Previewsection (outside any code fence) has a first non-empty line starting withevidence:and a reason. The template's<…>placeholder, left unedited, is not a reason.
A pass is journalled
evidence, with the plan's path, the commit and the sha. Any miss — no plan, a plan onmainbefore the branch, a section added in a later commit, an unapproved body — falls through to the approval path, unchanged. - the default branch is known (
The record has the trust of the approvals file: a workflow aid an agent
could write to, not proof. That is why every pass is journalled. The
exemption lapses when the screenshots show a visible change the plan did
not decide; that part is the skill's (worktree-task F7) to enforce: the
hook can only see that the person approved the declaration.
What counts as an interface change
The repository-level decision above says whether a repository is gated at all. Within a gated repository, one changed file counts only when all hold:
- The path. It is under
app/,src/orweb/(at any depth), or it is a.tsx,.jsx,.cssor.htmlfile. - The package. Its nearest
package.json— walking up from the file to the repository root, read from the pushed commit (git show <commit>:<path>), not the working tree — has adevscript, or listsreact,react-dom,react-native,react-strict-domor@duro-app/uiindependenciesor as a required peer dependency.devDependenciesand optional peers (peerDependenciesMeta.<name>.optional) do not count: duro-design-system'spackages/clicarries@duro-app/uiexactly there, and renders nothing. A monorepo whose root has adevscript is therefore judged package by package. Nopackage.jsonat all on the way up, or one that does not parse: the path decides. - Not comments only. For a
.ts,.tsx,.js,.jsxor.cssfile, the diff (git diff -U0 <base>..<commit> -- <file>; on a new branch with no base, each unpublished commit against its first parent) is read. When every added and removed line that is not blank is a comment line — starting with//,/*,*,*/, or a JSX{/* … */}— the file does not count. A line with code after a closed comment (/* a */ foo()), a*line that reads like code (CSS* {), a binary file or a diff that cannot be taken all count. Any doubt counts.
git config amont.agent.push-preview.ui still overrides the repository
decision; it does not change how files are judged.
A push to main or master is left to amont's branch-protect.
Supported push shapes
v1 reads an explicit remote plus explicit refspecs — git push -u origin feat/x, git push origin HEAD:feat/x, several refspecs — with -C <dir>
and a leading cd honoured. A destination under refs/tags/ is excluded;
one under refs/heads/ is judged, including a tag's commit pushed onto a
branch.
Everything else is unresolvable and journalled with its shape: a bare
git push, a remote with no refspec (git's push.default would decide),
--all, --mirror, --tags, deletes, glob refspecs, a URL as the remote,
a word from a substitution, an unknown flag. Under advise an unresolvable
push passes with a journal note; under deny a push to a UI repository is
held — --all is not the way around the gate. The shapes the journal
collects decide what the next version reads.
What push-published records
It never speaks. Before a push to a UI repository the hook remembers where
each branch destination stood (keyed by tool_use_id); after it, one
git ls-remote per destination, and one journal line:
| outcome | meaning |
|---|---|
published-approved | the remote now holds an approved UI commit it did not hold before |
published-unapproved | the same, without an approval — the gap advise leaves |
published-no-ui | published, nothing a person could preview |
already-present | the remote held the commit before the push |
unverified | the remote does not hold it, or could not be asked |
What this is not
A registration is the agent's attestation of worktree, commit and URL, and its guide is the agent's account of what it checked. Nothing here proves the browser served that commit or that the person looked. It covers pushes made through Claude Code, not pushes typed in a shell, and it is a workflow aid, not a Git enforcement boundary.
Reading the soak
grep -E 'push-(preview|published)' ~/.claude/amont-agent/journal.log
Every line names the repository the push ran in or the registration belongs to — never the session's working directory, which may be another worktree.
The rule's own lines (registered, replaced, approved, dropped,
unanswered, expired, prefilled, unbound, plan-approved,
evidence) carry the answer latency on
approvals and drops: option,184s,dur=0ms. The latency is the first
number (184s): the hook's own clock, from the PreToolUse of the marked
question to its PostToolUse. dur= is the payload's raw duration_ms,
kept only for comparison — it is not the person's answer time (a
minutes-long approval arrived as 0). A - in place of the seconds means
the question's PreToolUse record was written by an older release. A median
under about ten seconds after a week says the approval is a rubber stamp;
the evidence lines say how often a plan made the question unnecessary.
The review panel
Every plan passes a panel of expert reviewers before the person is asked to
approve it (ADR-0022, fleet rule work.plan-review-panel). The skill
/plan-review runs the panel. The plan-review-panel rule checks, at
ExitPlanMode, that it did.
Who reviews
The panel follows the areas of every repository the plan names:
| area | reviewers (plan-review-<role> agents) |
|---|---|
| always | language for each language, with lang=<l> in its review block, and backend |
| ops | platform |
| a command line | unix and tui |
| large (500 or more files or own commits) | po and architect |
| an interface | react, ui-design, ux-research and game-ux |
Areas are read from each repository's default tree, never from keywords in
the plan. The tree is the first of origin/HEAD, origin/main,
origin/master and HEAD that exists, with no fetch. Each signal:
- Languages:
Cargo.toml,go.mod,package.jsonortsconfig.json,pyproject.toml. - Interface: the
uitrait in.adr.yamlareas:, or apackage.jsonwith adevscript or a runtime React or@duro-app/uidependency. - Ops: a kustomization,
Chart.yaml, aHelmReleaseor FluxKustomization, or*.tf. - Command line: the
clitrait,src/main.rs,[[bin]], apackage.jsonbin, orpackage main.
Vendored, lock and dist/ paths do not count. A fork's commits are counted
past its upstream, and a shallow clone is judged on files alone.
What the plan carries
-
A
## Review panelsection, at most 5 lines, under the H1. -
The full reviews under a final
## Full reviews (reference). -
A machine comment as its last line:
<!-- panel: repos=amont-agent,decisions adds=ui reviewers=… body-sha=<12 hex> -->
repos= names each repository. A name resolves to
git config amont.agent.plan-review.<name>.path when that is set, for a
repository kept elsewhere; otherwise to $AMONT_AGENT_PLAN_ROOT/<name>, and
by default to ~/Developer/Perso/<name>. adds= declares areas the plan
creates: cli, ui, ops, large, or lang:<language>. A repository
that does not exist yet needs a lang: entry.
Which agents to launch
amont-agent plan-panel <plan.md> answers with the hook's own code:
repos=amont-agent areas=cli,lang:rust body=cf3f43c938c4 panel=full
plan-review-backend
plan-review-language --lang rust
plan-review-tui
plan-review-unix
panel=delta lists only the reviews needed since the plan's last accepted
body. panel=current lists none.
Binding a review to the plan
amont-agent plan-sha --block [--lang <l>] <plan.md> prints the review
block that each reviewer's prompt carries:
<<<PLAN path=/Users/me/.claude/plans/p.md sha=<64 hex>>>>
The sha is that of what the canonical body says: the plan without its
front matter, its review section, its full reviews and its machine comment,
read as CommonMark. Every word, number and operator counts, code blocks and
code spans count exactly as written, and so do link destinations and the
kind of each block (list item, quote, heading, table cell, an ordered
list's start). Blank lines, line wrapping, indentation outside code, table
padding, list-marker and emphasis style and escapes do not. A landed copy in
docs/plans/, which gains front matter and may pass through a formatter
such as prettier, therefore hashes like the approved plan, unless the
formatter rewrites code. Writing the review results into the plan never
makes a review stale. Editing the body does.
The Markdown parser is pulldown-cmark, pinned exactly: a new version may
read a plan differently, so a bump is its own change, and a fixture test
fails when it re-hashes.
Until the transition ends (2.25 onward), a review or a baseline bound
to the byte sha that 2.24 and earlier computed still counts. The hook's
journal line for a pass starts match=legacy when it rested on one, and
plan-sha --legacy prints that sha, for a plan whose machine comment was
written before. Both go once the journal shows no match=legacy for
14 days.
A delta in a new session. A plan's baseline is found by its path, by
its body, and, last, by its H1 and repos=: the same plan presented from a
new session has a new file name, and after an edit only its title still
names it. Two plans with the same H1 and repositories therefore share a
baseline; backend still reviews the body. If the sha cannot be computed
(the parser panicked), the hook asks the person instead of passing.
A review counts only when Claude Code recorded it as completed:
- in the foreground, a result with
status: completedand the agent'sagentType; - in the background, a task notification that Claude Code queued itself.
Notification text inside a tool result, or in a message, is never read.
The declared sha stops a skipped review and a review of an older body. It does not stop a model that deliberately declares a sha for a body it did not have reviewed.
What the hook answers
| situation | answer |
|---|---|
| first presentation | the whole panel, each role bound to this plan's path; backend bound to this body |
| body changed since the last pass | backend plus every area new since then, bound to this body |
| same body, no new area | passes (a re-presentation, or a metadata-only write) |
no machine comment, bad repos=, no review section | refused |
| no transcript, no plan path, git failed | the person is asked (ask) |
| refused twice already | the person is asked, with UNREVIEWED: refused N× |
⚠ unreviewed: … in the review section | the person is asked |
A pass is remembered under ~/.claude/amont-agent/plan-review/, which is
0700 with 0600 files, by plan path and by body. The same plan presented from
a new session, under a new file name, is therefore recognised.
The default stance is deny. Set amont.agent.plan-review-panel.stance to
advise to be told without being refused, or to observe to only journal.
The implementation review
Before the push of a branch that carries a plan, one independent reviewer
has read the diff against that plan and the repository's active rules, and
its verdict is recorded in the plan (ADR-0022, fleet rule
work.implementation-review). The worktree-task skill runs the review
(step F4b). The implementation-review rule checks, at git push, that it
happened for the tree being pushed.
The plan step has a panel; the quality loop has one reviewer. A diff with its surrounding code is larger than a plan and arrives on every push, and the reviewer's job is narrower: findings only, never an edit, never a waiver.
What the reviewer reads
Its prompt carries, in order: the review block, Round 1. or Delta., a
brief the session builds once, and its output contract. The brief is the
landed plan's path and Verification section, the diff against the default
branch without docs/plans/ (its stat when it is large), the repository's
active constraints (aval rules --level constraint), and the plan's
Non-goals. It checks plan conformance, that every recorded verification is
an observable check, violated constraints (rule id and file:line), tests
that cannot fail, and scope past the phases or the Non-goals.
Its result begins with the verdict, exactly, because the hook reads it:
Verdict: approve | approve-with-changes | rework
Findings:
1. [blocking|high|medium|low] <problem>. Evidence: <file:line>. Rule: <id or —>. Edit: <concrete change>.
Would still check by hand: <one to three things>
What a review binds to
The canonical tree id: the sha256 of git ls-tree -r of the commit
with every entry under docs/plans/ left out, 64 hex in any object format,
computed read-only. Recording the review in the plan therefore never makes
it stale; a code change does.
amont-agent tree-sha [--block] [-C <dir>] [--] [<rev>]
--block prints the block the reviewer's prompt carries:
<<<TREE repo=<name> sha=<64 hex>>>>
<name> is the basename of the directory holding the repository's common
.git, so a worktree checked out as amont-agent-wt-x says amont-agent.
Exit 0; 1 outside a repository or for a rev that names no tree (the reason
on stderr); 2 on a usage error.
What the hook does at a push
- Resolves the push (
crate::push_target): the repository and the commit each refspec publishes. A dry run, a push tomain/master, and a push of tags only are not its business. - No plan on the branch:
docs/plans/unchanged betweenmerge-base(<pushed>, <base>)and the pushed commit, where<base>is the first of<remote>/HEAD,<remote>/main,<remote>/masterthat exists, never the branch's own tracking ref and neverHEAD. The rule declines, silently. - Computes the canonical tree of the pushed commit.
- Looks for a binding: a completed
implementation-reviewagent in this session's transcript whose block names this repository and tree, or a pass remembered under~/.claude/amont-agent/implementation-review/by-tree/. Completion is read from structured fields only, as the review panel reads it; a notification echoed inside a tool result never counts. A reviewer resumed withSendMessageon a new block counts as a review of that block (below). - Reads the verdict. A review that said
reworkand got no later delta on the same tree is not a pass. - On the first transcript binding that passes, writes the pass file (0700 directory, 0600 file), so a new session pushing the same tree passes without a transcript.
| situation | advise (ships) | deny |
|---|---|---|
reviewed this tree, verdict not rework | silent | passes |
| no plan on the branch | silent | silent |
| reviewed an older tree | note, names both trees | held, names both trees |
verdict rework, no delta | note | held |
| never reviewed | note | held |
| transcript unreadable, git failed, no default branch known | note | held |
a push shape this guard cannot read (--all, --mirror, …) | journalled, passes | held |
Every note ends with the next step:
amont-agent/implementation-review: no implementation review of amont-agent tree 5085edae1d66 is in this session. worktree-task F4b: `amont-agent tree-sha --block`, then launch the implementation-review agent …
A resumed review
The delta (F4b.6) may resume the same reviewer with SendMessage, giving
it the new block, instead of launching a fresh one. Each round is a review
of its own, read from the shape a real session recorded
(tests/fixtures/sendmessage-resume.jsonl, 2026-10-08):
- the launch's result carries
toolUseResult.agentId; - a
SendMessagewhoseinput.tois that agent id, and whose result carriestoolUseResult.resumedAgentId= the same id, starts a round: its prompt isinput.message(the block is read from it) and its id is the SendMessage's own; - the round completes only on a structured task notification whose
<tool-use-id>is that SendMessage's id and whose<task-id>is the agent id. The launch's own result is never overwritten.
A message to an agent that was not launched as implementation-review, a
round whose notification names another task, or a notification typed or
echoed rather than queued by Claude Code does not count. Only this rule
reads resumes: the plan-review panel counts fresh launches alone.
The person's overrule
A rework that survives the delta goes to the person on a marked
question: its text starts with [implementation-review <repo>@<sha64>]
and its options are exactly Overrule, Fix, Hold. As with the preview
approval, the hook records the question before it runs and trusts the
answer only when that record exists and the question did not arrive with
answers pre-filled. On Overrule, and only when this session's
transcript shows a completed review of that tree that said rework, the
hook writes the pass file with verdict: overruled. An overrule that
bound nothing says so in the model's context.
Only the hook writes a pass file
Under deny a pass file is a trust boundary. A Write, Edit or MultiEdit
under ~/.claude/amont-agent/implementation-review/ is refused, with .
and .. folded before the comparison; a Bash command that would write
there — a redirect into it, or rm, cp, tee, python, chezmoi and
the like with a word naming it — is refused too, whatever the rule's
stance. Reading the record (ls, cat, find, grep) is allowed. The
Bash half is best-effort by nature: a parser cannot see inside python -c
beyond its words.
What the journal records
One line per push, under the rule's id, in the buckets amont-agent status counts: a declined push as unconfirmed with the reason
(reviewed, overruled, no-plan, dry-run, no-branch,
unresolvable), a spoken one as advised, denied or watched. The
excerpt of a spoken line is structured:
tree=<12 hex> review=<stale|rework|missing|unknown> verdict=<approve|approve-with-changes|rework|->
Away from a push: overruled, answered, unbound, unknown,
prefilled, unrecorded for the marked question, and denied for a
refused write. The soak reads them with:
grep -E ' implementation-review ' ~/.claude/amont-agent/journal.log
Stance
Ships advise. git config --global amont.agent.implementation-review.stance deny holds a push whose review is missing, stale, rework or unknown;
observe journals and says nothing. The 2026-10-06 soak review, alongside
push-preview, decides between deny, advise and retiring the gate from
the journal's review= counts and the first real cycles' finding quality.
The session notice
There is a mistake no command-level rule can catch, because no command is
wrong: a session opens in a checkout last pulled on Tuesday, the model reads
the tree it is given, and builds a feature that landed on origin/main on
Wednesday. The work is correct against the code it can see.
So at SessionStart the guard does the one thing the model cannot do for
itself.
The stale checkout
It refreshes origin/main — one branch, no tags, killed at five seconds,
skipped when FETCH_HEAD is under ten minutes old so a burst of sessions
shares one round-trip — and if HEAD is behind, says so:
amont-agent/stale-base: this checkout of amont-agent (branch main) is 8
commits behind origin/main; newest there: d3b2ed5 chore(release): 2.1.0
(3 days ago). Work that seems missing here may already exist on
origin/main — `git log HEAD..origin/main --oneline` lists it — and a
branch or worktree started from HEAD inherits the gap; one started from
origin/main does not.
It never pulls. git pull rewrites the working tree under whoever is
using it, and a per-task worktree exists precisely so that nobody does that.
Moving refs/remotes/origin/* is safe in every worktree at once; moving HEAD
is not.
If the fetch fails or times out, the notice is computed against the last successful fetch and says so. When the checkout is up to date, or it is not a repository, or there is no remote, it says nothing.
The stale-base rule is the same fact at the moment it is about to be
inherited: git worktree add, git checkout -b or git switch -c from HEAD
or a local branch, while that start point is behind. The remote form
(… -b feat/x origin/main) is the remedy and never fires.
git config --global amont.agent.stale-base.stance observe # measure, say nothing
git config --global amont.agent.fetch false # never touch the network
git config checkout.defaultRemote forgejo # measure against another remote
checkout.defaultRemote is git's own key for "which remote is the remote", and
a repository mid-migration — origin a mirror going stale, a second remote
carrying the truth — sets it once for both git and the guard. With two remotes
and no preference the guard says nothing rather than guess.
The stale guidance block
The same moment is when an agent reads AGENTS.md and believes it, and follows
it for the whole session. A block generated two releases ago can be wrong
before any command runs.
This one is entirely optional and entirely about
amont. If — and only if — this
repository carries an <!-- amont:start --> block and amont is on your
PATH, the guard asks amont the question amont already answers:
amont agents-md --check
and reports drift. Two file reads decide whether to spawn anything at all, so
a repository with no such block costs no process. No amont on PATH means
this says nothing — which is the right answer for anyone who does not use it.
It reads amont's stderr, not its exit code, and that is not fussiness:
agents-md --check exits 1 for a file it could not read as well as for a
drifted one, so the exit code alone would announce staleness for a permissions
error. A session-opening notice that cries wolf is worse than one that
occasionally says nothing.
git config --global amont.agent.agentsMdNotice false # silence this half
Open plans
If the session opens inside a repository with a docs/plans/ directory
(ADR-0022), one line per active plan: its title and its first unchecked
phase (- [ ] …). A pointer file — a plan whose canonical copy lives in
another repository — prints the canonical path and the phases this
repository carries. At most five lines; the rest are counted.
amont-agent/plans: 1 open plan in docs/plans/ (read one on demand; they are history, not instructions):
- Settings toggle (2026-09-28-settings.md) — next: Phase 2 — route
Closed plans (done, abandoned, imported) print nothing. A plan with no
status, or one outside that list, is reported: nothing else reads plan
front matter, aval check included. Nothing from a plan's body beyond those
two lines is loaded — an old plan is history, and loading it into every
session would compete with the instructions that are current.
Configuration
Every key is plain git config, readable and removable without this tool.
The kill switches
| key | default | effect |
|---|---|---|
$AMONT_AGENT_OFF | unset | any value switches the guard off for that shell |
amont.agent.enabled | true | switches it off everywhere |
The environment variable is checked first, before any rule runs, because
reading a git config key costs a process and the variable costs nothing.
amont.agent.enabled is consulted only once something has already fired.
Both take git's own boolean dialect — true/false, yes/no, on/off,
1/0, case-insensitively — because git parses them, not us. A value git
refuses warns once and falls back to the default rather than failing.
Stances
| key | takes |
|---|---|
amont.agent.stance | observe | advise | deny — every rule |
amont.agent.<rule>.stance | the same, for one rule |
Most specific wins. See stances.
The session notice
| key | default | effect |
|---|---|---|
amont.agent.fetch | true | may a session opening touch the network at all |
amont.agent.agentsMdNotice | true | may it mention a stale amont guidance block |
checkout.defaultRemote | — | git's own key; which remote is the remote |
What is deliberately not configurable
Which file any of these keys can be set in. --global, then --system,
and nothing else. Not a committed file, and not the --local/--worktree
config of the repository you happen to be standing in — see
stances. Write them with
git config --global; a bare git config inside a repository writes
--local, where this tool will not look.
Whether a failure is silent. It always is. See what it will not do.
Whether the journal can affect a decision. It cannot.
What it will not do
It never emits allow. That would short-circuit your own permission
prompt, so a guard approving everything it has no objection to would have
switched off the permission system it was installed beside. Silence is how it
says "no objection".
Every failure path is silence. An unreadable payload, an unknown event, a command it cannot parse, a rule that panics, a journal it cannot write — all of them exit 0 having written nothing.
A hook that fails toward refusing gets in the way of work you knew was
correct, and the fix people reach for at that moment is to delete it from
settings.json, which switches off every rule at once. One that fails toward
silence loses a single firing. That trade is the whole posture, and it is why
the hook payload is parsed with serde_json::Value and hand-written accessors
rather than a derived struct: a field that is missing or has changed type
becomes "no opinion", not a parse error somebody would be tempted to treat as
an opinion. What the hook writes back is emitted by a hand-rolled escaper
(src/json.rs), so the reading and the writing share no representation.
It does not judge what it cannot read. Heredocs without terminators,
unbalanced quotes, eval — all opaque, and opaque never fires.
The unit of that is the pipeline, not the line. | chains one command's
output into the next, so a stage we cannot read makes the whole run
unreadable. &&, || and ; do not: a command there is as independent of an
unreadable neighbour as of any other clause, and treating the whole line as
opaque cost every rule on it — measured over 33,774 real commands, 218 had a
readable pipeline thrown away. Those are judged now, and check names the run
it could not read beneath the verdict.
Two things keep that honest. eval, source and . run in THIS shell and
can move it, so one of them still hides the whole line — otherwise a later
confirm would resolve a path against a directory we can no longer vouch for.
And a finding from a partly-read command never refuses: it advises, whatever
stance the rule carries. Total opacity would have let the command run, so
blocking on half a reading is the worst outcome available.
It does not phone home. No telemetry, no update checks, no fetches — with
one exception, which is git fetch against your own remote for the
session notice, and amont.agent.fetch false switches
that off.
No repository can change a stance. Stances are read from --global and
--system git config only — never from a committed file, and never from the
.git/config of the repository the agent is standing in, which is a file that
agent could write. See stances.
The journal
Every firing is recorded at ~/.claude/amont-agent/journal.log, redacted, and
never transmitted anywhere.
It only counts. Nothing in it may participate in a decision — the rules read the command in front of them and nothing else. A guard whose verdict depended on its own history would be one you could not reason about from the command alone, and could not test from a corpus.
Shell analysis
src/analysis/ answers one question about a single Bash command, before it
runs: which network transfers it may make, and which it must, with the
evidence for each count. request-fanout is its first consumer. The analysis
decides nothing; the rules do.
Layers
| layer | module | job |
|---|---|---|
| frontend | analysis::frontend | source → crate-owned IR, spans on every node |
| semantics | analysis::interp, state, domain | abstract interpretation: state, control flow, exit statuses, counts |
| command models | analysis::models | a client's arguments → its explicit transfers |
| policy | rules::request_fanout | whether the result warrants advice, and the words |
Everything in analysis is pure. The dialect is an input; nothing reads the
environment, the filesystem or the network. The legacy lexer
(src/shell.rs) is untouched and still serves every other rule.
Inputs
- The command text.
- The dialect —
bash,zshorunknown. The hook reads it fromSHELL;check --dialectsets it; the backtester always usesunknown, because transcripts do not record the shell. Where bash and zsh differ, anunknowndialect gets the join of both readings — never one of them silently.
Dialect differences that are modelled: whether the last element of a
pipeline runs in the current shell (zsh yes; bash only under
shopt -s lastpipe); word-splitting of unquoted parameters (bash splits,
zsh does not); an unmatched glob in a URL-shaped word (bash passes it
through; zsh refuses the command); break/continue with no loop in the
current shell, as inside ( … ) (bash warns and carries on; zsh leaves).
Where the command does not say which seq runs, seq 5 1 is either nothing
(GNU) or five numbers (BSD, macOS): both are allowed.
tools/shell-oracle/ checks all of this against bash and zsh actually
running generated scripts with stubbed clients — see its Cargo.toml.
Assumptions
Every "established" count holds under these, and the rule's message lists them:
- the shell is not killed from outside;
- redirections succeed;
errexitandpipefailare off unless the command sets them (then they are modelled);- no aliases are defined;
- inherited environment variables are unknown;
- the command has no positional arguments of its own.
What is counted
Explicit command-line transfers: the requests a client's arguments ask
for. Redirects, authentication hops, a client's own retries and pagination,
and defaults from ~/.curlrc are named as possible extras and never counted.
A count is an interval: at least lower, at most upper, where upper is a
number, saturated (too large to represent), uncapped (positive evidence
that nothing bounds it: a loop following next-page links, --paginate,
wget -r, while true with no way out), or unknown (no evidence either
way). Zero annihilates: code that provably never runs makes no transfers,
whatever it would have done.
A finite upper bound on a loop comes only from the counter pattern: a
literal initial value, a comparison as the condition's final command, exactly
one unconditional +1 of the counter in the body, and no other write to it —
including from a function the body calls, or from anything unmodelled — and
no continue that could skip the increment.
Pacing
Each path through the interpreter carries how long it definitely slept
(sleep 30, sleep 2m; a backgrounded sleep does not count). A loop is
paced when every path back to its head — the end of the body, or a
continue — slept: a continue that skips the sleep, or a sleep behind a
condition, leaves it unpaced. A paced call site records the interval and its
burst: the transfers it makes per iteration, an inner loop or --retry
included. The innermost paced loop is the one recorded.
Paced and unpaced transfers never join across if branches: a poll in one
branch and a one-off request in the other are two different kinds of load.
request-fanout exempts a call site paced at least 10 seconds apart whose
burst is within its budget: a poll, not a burst, and foreground-poll's
business if it runs in the foreground.
The supported subset
Sequences, &&/||, pipelines and !, if/elif/else, case with ;;,
;& and ;;&, for, for ((…)), while, until, { }, ( ), (( )),
[[ ]], [ ]/test, function definitions in both forms, local/typeset,
break/continue with levels, return, exit, set -e,
set -o pipefail, shopt -s lastpipe; $( ) and backticks, $v, ${v},
positional parameters, $@/$*, the ${v:-x} family (not evaluated: its
value is unknown), every quoting form, unquoted brace expansion.
Unknown and Incomplete
- Unsupported constructs —
eval,source,trap,xargs,parallel, a command name built at run time, a program that runs another program — may do anything: every variable becomes unknown, the command may exit or never end, and the region is reported as activity the analysis could not follow. It never erases a known count elsewhere in the command. - Incomplete means a resource limit ended the analysis (nesting depth, IR size, interpreter steps, call depth, recursion). What was found before it is kept; nothing after it is known.
- Cardinalities —
{1..999999999},seq, curl URL ranges — are computed arithmetically and never expanded.
A rule built on the analysis never fires on an unknown count alone.