amont-agent

A guard that reads a shell command before Claude Code runs it, and can observe it, advise against it, or refuse it.

It exists because of a class of mistake no git hook can see, since the mistake is in the command string and never reaches a hook at all:

git push origin main 2>&1 | tail -5

A pipeline's exit status is the status of its last command. tail succeeds at tailing an error message, so a rejected push, a push killed by a timeout, and a push that never left the machine all report success — and the trimming discards the error text too, so the failure is silent in both channels.

amont-agent is a PreToolUse hook. It is a single binary, it is opt-in, and it phones nothing home.

The shape of it

Three stances, and the middle one is the point:

stanceeffect
observerecords the firing and says nothing at all
adviseputs the reason into the model's context; refuses nothing
denyrefuses the tool call, with the reason and the remedy

Of the 38 rules, three ship as deny, twenty-two as advise and thirteen as observe. A rule is promoted only once your own transcripts say it should be — see measuring and graduating — and amont-agent rules shows the stance in force on your machine.

Start here

amont-agent install          # prints the settings block, writes nothing
amont-agent install --write  # merges it into ~/.claude/settings.json
amont-agent doctor           # is it installed, runnable, and actually firing?
amont-agent status           # every rule, its stance, and what it has seen

Relationship to amont

amont is a git-hook manager by the same author. The two are independent: neither needs the other, they share no code, and cargo tree here shows serde and serde_json and nothing else.

They meet in exactly one place, and it is optional. If amont is installed and this repository carries its generated AGENTS.md block, a session opening on a stale block is told so — see the session notice. With no amont on PATH, that check says nothing.

Relationship to attest

attest is the CI end of the same story: amont signs a note at pre-push naming the gates that really ran, and attest verifies it so CI can skip them. It has no connection to this guard beyond the author and the conviction all three share — trust what was verified, never what was reported. The masked push at the top of this page is what that looks like when it fails.

Installing

The binary first, then the wiring. They are separate acts on purpose: a program that can refuse the commands your agent runs should not install itself into your settings.json as a side effect of you fetching it.

The binary

curl -fsSL https://raw.githubusercontent.com/fredericrous/amont-agent/main/install/install.sh | sh
irm https://raw.githubusercontent.com/fredericrous/amont-agent/main/install/install.ps1 | iex

Either one downloads a release binary, verifies it against the published SHA256SUMS, and puts it in ~/.local/bin. Neither wires anything in.

Or: brew install fredericrous/tap/amont-agent, cargo install amont-agent, or a binary straight from Releases — Linux x86_64 (gnu and static musl) and aarch64 (gnu), macOS (Intel and Apple silicon), and Windows x86_64.

Why there is no npm package

There was one, briefly, and removing it is the more useful answer than listing install methods.

install bakes the ABSOLUTE path of the running binary into settings.json, deliberately: a PATH-resolved command exits 127 into Claude Code's non-blocking bucket the moment PATH differs, which disables the guard with nothing to notice it. npm cannot supply a stable absolute path. Under npx the binary lives in npm's _npx cache, which npm garbage- collects; under a project-local npm i -D it lives in one project's node_modules, while the guard it configures is machine-global — so rm -rf node_modules in one repository would silently disable the guard for every session on the machine.

amont has an npm package for a reason that does not transfer: it is a per-repository tool, and npm i -D amont plus a prepare script means the hooks travel with the repository. This is per-developer machine configuration — ~/.claude/settings.json, global git config, a journal in ~/.claude/. Nothing about it belongs to a project.

Every channel above hands install a path that stays put.

The wiring

amont-agent install          # prints the settings block, writes nothing
amont-agent install --write  # merges it into ~/.claude/settings.json

install refuses to guess. It will not patch a settings.json it could not parse, and it will not write one whose formatting it cannot reproduce — a diff full of reformatting hides the one line it added — so it prints the block and changes nothing unless --reformat says otherwise. uninstall removes exactly what it wrote and leaves everything else byte-identical.

amont-agent install --write --project   # .claude/settings.json instead
amont-agent install --write --local     # .claude/settings.local.json

Four kinds of entry are written: the guard on PreToolUse, for Bash and for the file tools; the assertions on PostToolUse for Bash, which check what a command that reported success actually did; a PostToolUse entry for Read, which remembers a file only once it has actually been read; and a SessionStart entry that leaves a heartbeat and states where the checkout stands against the remote. Without the heartbeat, doctor cannot tell "nothing fired this week" from "the guard has been dead since Tuesday".

Your settings.json keeps the mode it already had, and a file created here starts at 0600: it can hold MCP environment blocks, and those hold credentials.

Knowing it is alive

Claude Code hooks fail open quietly: a command that cannot be resolved exits 127, which is a non-blocking status, and nothing tells you. doctor exits non-zero when the guard is inert, so it can run from cron:

✓ installed in /Users/you/.claude/settings.json
✓ amont-agent 2.0.0 at /Users/you/.local/bin/amont-agent
✓ a refused command produces a valid decision document
✓ last ran 4m ago
✓ acting on pipe-to-tail

Turning it off

AMONT_AGENT_OFF=1                                # this shell only
git config --global amont.agent.enabled false    # everywhere
amont-agent uninstall --write                    # remove the settings entries

Demoting one rule is almost always the better move than switching the guard off — see stances. A guard that is hard to back out of is one people uninstall instead of demoting, and uninstalling takes every rule with it.

Stances

A rule does one of three things, and the middle one is the point:

stanceeffect
observerecords the firing and says nothing at all
adviseputs the reason into the model's context; refuses nothing
denyrefuses the tool call, with the reason and the remedy

observe and advise are not two ways of saying "not blocking yet". additionalContext enters the model's context and therefore changes its behaviour, which contaminates the rate the observation exists to measure. A rule that talks is intervening. That is why the two are named differently and why backtest numbers from an advise rule are not comparable with the numbers that justified promoting it.

That contamination is also what makes it measurable. amont-agent backtest --compliance reads the two stances against each other: what a model did after an advise fired, beside what it did after an observe fired and said nothing. The second is the rate the habit corrects on its own, and an advise that does not beat it is costing context tokens for nothing. See measuring and graduating.

Changing one

Takes effect on the next command; nothing to restart.

git config --global amont.agent.pipe-to-tail.stance observe
git config --global amont.agent.stance observe          # every rule

The ladder, most specific first:

rule.default_stance  <  amont.agent.stance  <  amont.agent.<id>.stance

then capped at the rule's own ceiling, and clamped to observe if the guard is switched off.

Ceilings

Every rule declares the loudest stance it may ever take. For most rules that is deny, and the ceiling changes nothing. A rule whose finding is an estimate — "this loop may make hundreds of requests" — is capped at advise: refusing a command on an estimate would claim a certainty the analysis does not have.

The cap is applied after every configured key, so neither amont.agent.stance deny nor the rule's own key can pass it, and graduate --to deny refuses a capped rule with the reason. amont-agent rules prints each rule's ceiling beside its stance.

Why git config and not a committed file

Promotion power stays on the machine, with the person. A rule that a committed file could promote to deny would mean cloning a repository hands it the power to refuse your shell commands.

This is enforced three times, because neither of the first two was enough.

  1. amont.conf's parser will not name these rules, so a committed manifest cannot reach them.
  2. amont's config reader lets a repository's committed policy set lines outrank system and global git config — right for a hook manager, wrong for a guard — so this project reads git config itself, without that ladder.
  3. Git's own search order ends at --local and --worktree, and the last file wins. A .git/config therefore outranked your --global answer without any of amont's machinery being involved — and the agent whose command was just refused can write one, since git config is not a command any rule here objects to. So the reader takes --global, then --system, and nothing else.

A stance answers to your own git config and to nothing a repository carries or a process standing in one can write. A file your global config includes or includeIfs counts as your own: git skips those for a scoped read unless asked, and this reader asks. graduate and demote write --global for the same reason.

One consequence worth stating plainly: git config amont.agent.<rule>.stance run inside a repository writes --local by default, and this tool will not read it. Pass --global, which is what every example here does.

The rules

ruleships aswhat it catches
pipe-to-taildenya mutating command whose status is swallowed by a pipe
bare-stash-popobservegit stash pop with no ref, where refs/stash is shared across worktrees
gh-pr-merge-autoobserve--auto on a repository with no required checks, which merges immediately
publish-without-skilladvisea v* tag pushed or a pull request merged with no call of the tag-release / merge-when-green skill in this turn or the one before, so the procedure ran from memory
forge-merge-by-handadvisemerging a pull request by POSTing to the forge's merge endpoint, which answers 200 whether the checks passed, failed or never started
forge-status-stale-rowobservekeeping the first row of a commit's append-only /statuses list, so a stale pending reads as the present and the wait never ends
no-verifyobserveturning the whole commit gate off rather than one check
git-add-broadobservestaging the tree instead of the change
stale-baseadvisea branch or worktree started from a checkout the remote has moved past
push-preflightadvisea git push whose slow pre-push test gate has not been rehearsed with amont rehearse --wait
push-previewadvisea push that would publish interface changes no approved localhost preview covers, unless the plan the person approved declares the evidence is enough (preview approval)
plan-review-paneldenya plan presented at ExitPlanMode before its expert review panel ran (the review panel)
plan-phases-opendenya turn ending while the plan this branch carries still has an open phase that is not a 🧑 decision: — the agent is sent on to it
implementation-reviewadvisea push of a branch that carries a plan, whose diff no independent reviewer has read for the tree being pushed (the implementation review)
foreground-polladvisea polling loop or gh run watch in the foreground, where the tool's ten-minute clock will kill it one poll short
unbounded-background-pushadvisea git push sent to the background with its output kept in a file and no timeout, so the pre-push gate inside it can run for as long as it likes with nothing ending it or reporting it
sed-in-placeadvisesed -i spelled for the other sed (-i '' on GNU, bare -i on BSD)
kubectl-gitopsadvisean imperative kubectl write in a repository Flux or Argo reconciles
tag-after-commitadvisegit tag chained onto a git commit that a hook may have refused
release-tag-pushobservepushing a v* tag, which publishes: the tag may name the wrong commit, and a green workflow does not mean the artefact is right
worktree-remove-forceadvisegit worktree remove --force on a worktree that still holds uncommitted work
amend-pushedadvisegit commit --amend on a commit the remote already has
branch-force-deleteobservegit branch -D on a branch whose commits are on no remote and not merged
poll-blank-verdictobservea wait that stops on any value but the one it names, so a failed lookup's empty string reads as the answer
worktree-isolationobservea branch created, or git reset --hard, in the primary checkout of a repository that already has linked worktrees
stdin-hangobservea command that will read standard input with nothing on it — cat > file, a bare interpreter, tee outside a pipe — which blocks silently until the tool's clock runs out
glob-in-flag-valueadvisean unquoted glob inside a flag value (--include=*.ts), which zsh expands — or fails on — before the program sees it
glob-no-matchadvisean unquoted glob operand that matches nothing, which under zsh aborts the clause before it starts while a later clause reports success
equals-separatorobservea bare word beginning with = (echo ===, [ x == y ]), which zsh reads as a command lookup and whose failure aborts the whole command list
unsplit-expansionobservean unquoted $var under set -- or for … in, which zsh leaves as one word: the positionals or the loop get the whole string and the command runs wrong while reporting success
path-operand-missingadvisea read of a path that is not there — under the grep shim a warning in mid-stream, for the coreutils one line and carry on — which the chain then reports as success
stat-bsd-formatadvisestat -f '%…' on GNU stat (or -c on BSD), which prints a filesystem report where a timestamp was wanted
whole-file-dumpadvisea file poured whole into the tool result by cat/sed -n/head — 31% of all result bytes measured — where the Read tool would have windowed it
persisted-output-dumpadvisereading back whole a tool result the harness saved to a file for being too large, paying for it twice
file-rereadadvisea Read, or a cat, of a file this session already has in context and that is unchanged on disk since — answered from the session's own record, not the command
read-unbounded-largeadvise (ceiling deny)a Read with no offset/limit of a file over 16 KB, where the whole file lands in the context and every later turn carries it; plan files and diffs a reviewer is handed are exempt
lint-suppression-addedadvise (ceiling deny)an Edit, MultiEdit or Write that adds a lint suppression (# type: ignore, # noqa, eslint-disable, @ts-ignore, #[allow], //nolint) or loosens a lint configuration (tsconfig, eslint, pyright, ruff, Cargo [lints], golangci), where general.no-disabled-safety says to fix the finding instead; the file is compared before and after by marker, not by line
request-fanoutobserve (ceiling advise)one command that may make more than 50 explicit network transfers to one destination — a loop following Link: next, gh api --paginate, a curl URL range — counted by the shell analysis (analysis.md) with the loops and calls that multiply them

Two more checks run after a command rather than before it, on PostToolUse: push-landed and push-published verify what a push that reported success actually did. See the assertions.

amont-agent rules prints this with each rule's measured firing rate, and with the stance in force on this machine rather than the shipped one: a rule you promoted in ~/.gitconfig reads deny (ships as observe), and a rule with a ceiling names it, (max advise).

Why only three of them deny

pipe-to-tail blocks because seven consecutive weeks of measurement showed no downward trend while every other habit halved. That is the bar: a rule earns deny from your own transcripts, not from an argument about how bad the mistake is. A habit the model is already correcting does not need a deny. See measuring and graduating.

plan-review-panel is the other kind of deny, and the person's choice rather than a measurement: it fires on ExitPlanMode, not on a shell command, and checks a fact — whether the review panel the plan calls for ran on the plan being presented — naming the missing roles when it did not. After two refusals of the same plan, or whenever it cannot check, it hands the call to the person as ask instead. See the review panel.

plan-phases-open is the third, also the person's choice. It fires on Stop, the moment the agent ends its turn. ADR-0022 makes a plan one branch and one pull request with its phases as commits, so a turn that ends with a phase still open is a stall, not a handoff. The rule reads only the active plans changed on the current branch (against the merge-base with the remote default branch), never one already on main, and sends the agent on to the first open - [ ] under ## Phases. It is silent when that phase is a 🧑 decision:, in plan mode, while background tasks are still running, and when the agent's last message has a line starting with WAITING: <reason> (a preview awaiting approval, a deferred phase, a blocker). After three continuations on the same phase in one session it lets the turn end and tells the person, with a note only they see; their next prompt resets the count. Anything it cannot establish — no repository, a git failure, an unreadable plan, a session id it will not use as a file name — lets the turn end. To turn it off without uninstalling: git config --global amont.agent.plan-phases-open.stance observe.

stale-base advises from the start because it refuses nothing, speaks only after measuring a real gap, and names a failure no correcting loop can see — nothing fails when you build on stale code. The work is correct against the code it can see, and the conflict arrives later, from somewhere else. So a session opening in a checkout the remote has moved past is told — see the session notice for why it fetches and never pulls.

push-preflight advises for the same reason. git opens its connection to the remote before it runs pre-push and holds it idle for as long as the test gate takes; a remote that closes idle sessions kills the push after the gate has already passed, and the model reads "the network" where the cause was the gate's placement. With amont ≥ 1.28, amont rehearse --wait runs the same gate on a snapshot of HEAD with no connection open — or follows the rehearsal amont.rehearseOnCommit already started — and stamps the tree, so the push that follows skips the suite (amont run pre-push on 1.27). The rule speaks only when confirm finds all three facts: amont guards this repository's pushes, a test gate would run for this push, and HEAD's tree carries no stamp yet.

Shape, then the world

Most rules after the first handful came out of the transcripts the same way — tens of thousands of Bash calls, sorted by what failed, was killed, or drew a correction (see mine). Each fires on shape and, where the fact lives in the world, confirms it first: whether the call already runs in the background, which sed is on PATH, whether the repository holds a Flux or Argo resource, whether the worktree is dirty, whether the remote has the commit, whether any other branch has the commits. A rule that fires on shape alone, like tag-after-commit, names a failure every command in the chain reports as success.

pipe-to-tail in full

git push origin main 2>&1 | tail -5

A pipeline's exit status is its last command's. tail succeeds at tailing an error message, so a rejected push, a push killed by a timeout, and a push that never left the machine all report success — and the trimming discards the error text, so the failure is silent in both channels.

The remedy the rule prints is not "don't use tail": it is to run the mutating command on its own, read its output afterwards, and then verify the effect (git ls-remote origin refs/heads/<branch>) rather than the exit code.

set -o pipefail first, or writing to a file and tailing the file, are also accepted — the rule fires on the shape, and the reason names all three ways out.

Asking about one command

amont-agent check 'git push | tail -1'

No stdin, no session, no journal entry — just the rules over one string, with whatever they would have said.

The assertions

A rule reads a command before it runs. An assertion checks what a command that already reported success actually did.

The two halves are not the same job. pipe-to-tail can refuse git push … | tail -5 because the mistake is visible in the command string. Nothing in a command string tells you that the push you just ran reported Everything up-to-date about a branch you were not on.

idfires onasks
push-landedgit pushis the branch on the remote, at the commit you have?
push-publishedgit push to a UI repositorydid it publish a new commit, and was it an approved preview? (records only; see preview approval)

Only successful calls

Claude Code sends a failed tool call to PostToolUseFailure, an event this crate ignores. So everything an assertion sees claimed to work — which is the whole point. A command that returned an error is already in front of the model; there is nothing invisible left to point out.

An assertion cannot refuse

The tool has already run. A deny stance therefore speaks exactly like advise, the same way it does at a session opening.

What an assertion can do is state a fact:

amont-agent/push-landed: `git push` exited 0, but origin/main is at d0765d4f9
while the local branch is at 8970c7ffb. The push did not land. Run the push
again on its own and read its output, then confirm with `git ls-remote`.

That is not advice to weigh. It is the remote's answer.

What it refuses to judge

push-landed handles the unambiguous shapes and nothing else. A HEAD:refs/heads/other refspec, a tag push, several refspecs at once, --delete, --mirror, --all: each needs a different question asked of the remote, and a confidently wrong accusation costs the channel its credibility — which is the only thing the channel has.

The same goes for everything it cannot establish: a detached HEAD, a remote given as a URL, a working directory that has gone, a remote that cannot be reached without a password. All silence.

It can never prompt

Hooks run with no controlling terminal, so a credential prompt does not fail — it hangs, and it hangs the session rather than this process. Every child runs with GIT_TERMINAL_PROMPT=0, GIT_ASKPASS/SSH_ASKPASS disabled, ssh in BatchMode, and a five-second deadline. A guard installed to make pushing safer must never be the reason a push becomes impossible.

Measuring

examine is pure, so the backtester replays it like any rule:

$ amont-agent backtest push-landed --since 2026-07-06
push-landed   1007   38.7   routine   observe

Read that number correctly. For a rule, a firing is a mistake caught. For an assertion, a firing is a question asked — 1,007 pushes in 29,758 Bash calls, about one call in twenty-six paying for one git ls-remote. It speaks only when the answer disagrees, which is far rarer.

How often a claim is actually broken is not backtestable at all: verify touches the world, and the world has moved since those commands ran. That number accumulates forward, in the journal:

$ grep push-landed ~/.claude/amont-agent/journal.log

held is a push that landed, broken one that did not, unverified one this crate declined to judge.

Measuring and graduating

This is the part that makes the rest defensible. Every rule's stance is a claim about your own behaviour, and the claim is checked against your own transcripts rather than asserted.

The loop is: mine → backtest → explain → review → compliance → graduate.

0. Mine — what is going wrong that no rule names?

Every step below prices a rule somebody already thought of. mine is the step before that: it groups the transcripts by command shape and ranks the shapes that went wrong.

amont-agent mine --since 2026-08-15
amont-agent mine --min-support 10 --min-rate 0.4
amont-agent mine --format cases >> tests/corpus/<new-rule>.cases

A shape is the parsed command with its literals masked — the branch, the path, the line number and the commit message taken out, the program, the verbs and the flag set kept:

git push origin feat/mine 2>&1 | tail -5  ─┐
git push origin fix/lint  2>&1 | tail -20 ─┴→ git push origin <word> 2>&1 | tail -5

A call counts as gone wrong when either of two things the transcript records is true:

  • failed — the tool result carried is_error: a non-zero exit, a refused permission, a run the harness killed;
  • corrected — a near-identical command followed within a few tool calls. A model that re-issues the same shape is a model whose first attempt did not land, and that is the interesting half: the mistakes worth a rule are the ones nothing reports.

Only the HEAD of a run of near-identical calls counts as a correction. A polling loop is one decision repeated twenty times, not nineteen mistakes — before that rule existed, the top line of the real report was echo waiting.

bad and rate are suspicion, not verdict. A shape re-run for good reasons carries a high rate and names no mistake at all. That is why every row prints its samples and why the next step is a person reading them.

Shapes an existing rule already fires on are listed apart, covered by <rule>, rather than proposed — and a covered shape that is still going wrong is a rule that is observing when it should be advising.

What mine does not do is write the rule. Nothing learned here reaches the hook: the path is mine → write the rule by hand → cases → corpus check → graduate, and what ships is the hand-written rule with its reason, its remedy and its reviewed corpus. A guard that refused a command because a clustering run found it suspicious could not explain itself to the person whose work it just refused.

1. Backtest — what would this have cost me?

backtest replays your Claude Code transcripts through the rules and reports firings per 1,000 tool calls per week, so a rule's cost is a number rather than an impression.

amont-agent backtest --since 2026-07-06
amont-agent backtest --rule pipe-to-tail --json
amont-agent backtest --transcripts ~/.claude/projects   # where they live

A weekly series is the thing to read, not a total. A habit that is halving on its own does not need a deny; the model is already correcting. A flat line over weeks is a habit that will not correct itself, and that is what promotion is for.

2. Explain — look at the actual matches

A rate is only trustworthy if the matches behind it are real. explain prints every match for one rule so you can read them.

amont-agent explain pipe-to-tail
amont-agent explain pipe-to-tail --sample 20
amont-agent explain pipe-to-tail --sample 20 --rank novelty

--rank novelty changes WHICH twenty. The default is the first twenty the walk met — the oldest project, the oldest session, and, because a habit repeats, very often twenty spellings of one command. Novelty picks the twenty least like each other and least like the cases already in tests/corpus/<rule>.cases, by greedy max-min over the shape distance, so an hour of labelling buys as much of the precision estimate as an hour can. It is deterministic: the same transcripts and the same corpus pick the same cases, and a second pass does not hand back the first pass's.

3. Review — turn matches into reviewed judgements

Precision is kept as a corpus of judgements, not as a metric, because a metric charts a regression and a test prevents one.

amont-agent explain pipe-to-tail --format cases >> tests/corpus/pipe-to-tail.cases
$EDITOR tests/corpus/pipe-to-tail.cases    # each `?` becomes match or nomatch
amont-agent corpus check                   # and this runs in the test suite

Include the cases that should not match. A corpus of positives alone measures recall and says nothing about how often the rule is wrong, which is the number that decides whether it can be allowed to refuse anything.

4. Compliance — is the advice worth its tokens?

advise buys its place in the model's context with tokens, every session, forever. The backtest says what that costs. This says what it buys.

amont-agent backtest --compliance
amont-agent backtest --compliance --rule glob-in-flag-value --json
amont-agent backtest --compliance --window 40

For every firing it looks at what the model did next. If the next equivalent command — same shape, within the window — no longer matches the rule, the habit changed (complied). If it matches again, the advice was read and ignored, or never reached the model (ignored). If nothing equivalent followed, the firing is unanswered and counts towards neither.

Three things make the number readable:

  • Per model. Models differ, and a pooled figure hides which one is listening. claude-opus-5 and claude-fable-5 do not answer the same way to the same sentence.
  • observe rules are the control. They said nothing, so their share is the rate the habit corrects on its own. Advice is worth its tokens only where the advised share beats the silent one — the same argument the stance ladder rests on, measured instead of assumed.
  • deny rules are left out. A refused command never ran, so there is no next command to compare it against.

What it cannot see: a transcript records what the model ran, not whether the hook spoke. A firing here means "the rule as it stands today would fire on this command", replayed. Where a rule has been widened since, or was observing then and advises now, the number is a reconstruction — and it is still the only evidence there is.

The Read tier from transcripts

The Read tool never reaches the backtester, so read-unbounded-large is measured from the transcripts directly:

tools/read-rate.py
tools/read-rate.py --weeks 8

For each ISO week it prints the tool calls, the Reads over 16 KB with no offset/limit after the same exemptions the rule makes (plans, diffs, media, tool-results/), that count per 1000 calls, and the characters. Run it to recount the rule's per_1000 and again two weeks after a release, to compare against the weeks before.

The write tier from transcripts

lint-suppression-added never reaches the backtester either: an Edit or Write is not a command. It is measured from the transcripts with tools/suppression-rate.py:

tools/suppression-rate.py
tools/suppression-rate.py --weeks 8
tools/suppression-rate.py --until 2026-10-06

For each ISO week it prints the tool calls, the Edit, MultiEdit and Write calls that add a suppression or loosen a lint setting, and that count per 1000 calls. Run it to recount the rule's per_1000 and again two weeks after a release.

The script sees fragments only: a transcript holds what the model sent, never the file, so it compares old_string with new_string, as the hook's fragment fallback does. With no section around the text it undercounts the config rows that need one ([tool.pyright], [tool.ruff*], [lints.*], .golangci.yml disable: and exclude*), unless the edit carries the header itself. --self-test classifies tests/fixtures/suppressions.txt, which a Rust unit test classifies too; tests/suppression_rate.rs runs it.

What came before a publish, from transcripts

publish-without-skill asks a question the backtester cannot: whether the tag-release or merge-when-green skill was called in this turn or the one before. Its confirm reads the transcript, and the backtester never runs confirm. It is measured with tools/skill-rate.py:

tools/skill-rate.py
tools/skill-rate.py --since 2026-10-07 --list

It prints the publishing Bash calls per 1000 (a v* tag pushed, gh pr merge, a merge sent to /pulls/<n>/merge), split by what preceded them: no-skill, stale (called two or more prompts back), allowance (covered only by the one prompt the window allows) and covered. It prints the same counts per unique command per session, since a refused command is retried. Then it prints how old the covering call was, and the uncovered rate by week. The matcher is a regular expression. Against the Rust lexer it overcounts by about 6% (1,551 against 1,461 on 2026-10-07), so backtest --rule publish-without-skill remains the count of matches. Run it again two weeks after a release.

5. Graduate — promote on the evidence

amont-agent graduate bare-stash-pop --to advise
amont-agent graduate bare-stash-pop --to deny

Promotion is gated on the corpus: a rule cannot be promoted past a corpus that does not support it.

Demotion is not gated at all

amont-agent demote bare-stash-pop

No questions, no evidence required, effective on the next command. This asymmetry is deliberate. A guard that is hard to back out of is one people uninstall instead of demoting — and uninstalling takes every rule with it, including the ones that were working.

Preview approval

push-preview and push-published carry ADR-0028 (work.preview-approval, fleet decision corpus, replacing ADR-0023): a commit that changes a user interface is published only after the person approved that commit, seen on localhost — unless the plan the person approved declares the evidence is enough (below, rule work.preview-unless-planned-evidence).

The rule exists because interface regressions kept reaching pull requests and deployed sites after the agent had checked its own screen. The person's eyes are the check that does not share the agent's blind spots, and localhost is where rejecting a screen costs one sentence instead of a fix pull request and a release.

The flow

The person works on several projects at once and does not remember where this one stood. A bare "[preview f0223d3] Ship website-builder@f0223d3?" gives them nothing to decide with, so a preview reaches them as a guide (fleet rule work.preview-is-guided), three ways at once: the brief in the terminal, a local page, and the app already open in their browser.

  1. Verify. The agent drives the changed route in a real browser, checks the console and network, and takes before and after screenshots.

  2. Write the guide — markdown, outside the worktree (~/.claude/amont-agent/attestations/<sha>/guide.md by convention, the screenshots beside it), because a file inside it, or the page rendered beside it, would dirty the commit it describes. The format is below.

  3. Serve and register. From the clean worktree, at the commit to be pushed:

    amont-agent preview register --url http://localhost:5173/settings \
      --guide ~/.claude/amont-agent/attestations/abc1234/guide.md
    

    Run it in the foreground as the last command of its line. Anything may come before it, joined by && or ;:

    cd ~/Developer/app-wt-settings && npm run build && amont-agent preview register --url … --guide …
    

    It refuses a dirty tree, a non-commit HEAD, a guide inside the worktree and a guide missing a required section (exit 1; stderr lists exactly what is missing). Otherwise it renders the guide to index.html beside it and prints one JSON object, compact on one line:

    fieldwhat it is
    idthe registration; it starts with the commit's sha7
    repo, committhe checkout's top level and its full HEAD
    urlwhere the clean worktree is served
    guide, page, page_urlthe guide, its rendered index.html (absolute), and that page as a file:// URL
    label, aliasesthe names a person knows the commit by: the checkout's directory, then the main worktree's directory and the origin repository's name, each @sha7
    question_prefix[preview <id>] <label>: what the marked question starts with, verbatim
    attestationguide again, for hooks older than 2.21.0

    --open also hands the page to the platform opener (open, xdg-open, cmd /c start) without waiting; failing to open never fails the command.

    The PostToolUse hook binds that output to the session and to the person's current prompt, in the repository the command ran in — the leading cds followed — which must still be the printed repository at the printed HEAD; the journal line names the page. The JSON is read from the last non-empty line of stdout, so what earlier clauses print does not matter. The register must be the last clause: one that is followed by anything (register && echo), piped (| tee), redirected, inside $(…), after ||, or in the background is not bound, and the journal says why.

    A line an earlier clause printed cannot pass for the register's own (printf '<json>'; amont-agent preview register --bad): the PreToolUse hook stamps when the call started, in previews/registers/<tool_use_id>, and the bind requires that the id starts with the commit's sha7 and that the page exists, names the full commit and was written after that stamp. A call with no stamp is not bound. Stamps nobody came back for are swept after a day.

  4. Open the app and the page. Before asking, the agent opens the app in the person's browser at the state the first step reaches (the panel already open, the form already filled) and the rendered page beside it. That is the agent's job — the worktree-task skill's F7, with Claude in Chrome — not this binary's: register only renders the page and, with --open, opens it.

  5. Ask, in the same turn. The brief first — the guide's sections, in the terminal — then one AskUserQuestion whose text starts with the register's question_prefix ([preview <id>] <label>), with options exactly Approve, Request changes, Hold — no "(Recommended)" suffix.

  6. Push after Approve. The approval covers that commit; a new commit, amend or rebase needs a new preview.

--attestation <file> is accepted for 2.21.0 only, as a deprecated alias of --guide: the file is read as the guide and refused unless it is one.

The guide

Markdown with these five H2 sections. Headings match with emoji, punctuation and case aside (## 📍 Where we are: is Where we are); the page shows them in this order, then any other section the guide adds.

sectionwhat it carrieschecked
Where we areproject, branch and plan, and why the person is asked. Its first line is the page's titlepresent, not empty
What you should seein their words; "nothing new" when that is the pointpresent, not empty
Try itnumbered steps: the exact URL, each click named by its visible label and position, and after each step what should appearat least one numbered step (1.) and one http(s) URL
Referencebefore and after screenshotspresent, not empty
Already checkedwhat the agent verified, and what to look at especiallypresent, not empty

A full example (guide.md, with before.png, after.png and panel.png beside it):

# Settings: Save moves to the header

## 📍 Where we are
duro-app, branch `feat/settings-save`, plan **Settings panel** (phase 2 of 3).
You are asked because the Save button moved; nothing else on the page changed.

## What you should see
The **Save** button now sits in the header, top right, instead of at the
bottom of the form. Saving works exactly as before.

## 👉 Try it
1. Open http://localhost:5173/settings
   - You should see the Settings form, with **Save** in the header, top right.
2. Change **Display name** (first field) to anything.
   - **Save** turns from grey to blue.
3. Click **Save**, top right.
   - A green toast "Saved" appears bottom left, and **Save** is grey again.

![the header after step 1](panel.png)

## Reference
![before](before.png)
![after](after.png)

## Already checked
- `/settings` at 1280 and 390 px wide: no console error, no failed request.
- Keyboard: Tab reaches **Save** right after the page title.
- Look especially at the narrow window: the button must not cover the title.

The page is one self-contained file: inline CSS readable in light and dark (prefers-color-scheme), no script, no external asset. Its title is repo@shortsha — <first line of Where we are>; under it, a large Open the app link to --url, then the sections. Images (![alt](file.png), relative to the guide's directory) render inline; a missing one is a warning on stderr, not a refusal. When one image's file name says before and another's says after, the two show side by side (class="pair") at the top of Reference. The converter is hand-rolled for this subset — headings, paragraphs, bullet and numbered lists with one nested level, fenced code, quotes, bold, code, links, images and bare URLs — and escapes everything else; a link whose scheme is not http(s), mailto or file is text. No markdown crate: amont-agent is on a trust path and a page only the person reads does not clear change.dependency-bar.

Mockup mode

When the previewed branch commits a picked mockup — a *.dc.html artboard under docs/mockups/<screen>/ (or the directory git config amont.agent.preview.mockups names; the application-landscape ui-handoff PR check reads the same key) — the guide must prove fidelity (handoff.prove-fidelity, ADR-0016). Added to the five sections:

  • under ## Reference: the screen's path, a Viewport: 1120px · light line (the width and theme both images were taken at), and the picked artboard beside the built screen as image files next to the guide: artboard.png and after.png (mockup and live are read the same; states pair by suffix, artboard-empty.png with after-empty.png);
  • a section ## Differences from the mockup: None, or one bullet per difference, each - fixed: … or - deliberate: <reason>.

The range is the branch's own commits, from its merge base with the default branch — never from its upstream, which after a first push already holds the artboards. Artboards on a branch whose base cannot be resolved are a refusal, not a skip. A later branch that only edits a screen whose artboards are already on main is not in mockup mode; the PR check asks it for the side-by-side or Mockup: none — <reason>.

The images are copies of the proof PNGs committed next to the artboards, so copying them beside the guide comes first, as its own command: a register chained after a cp is not bound (see below). Two PNGs whose widths differ by more than a tenth are a warning. On the page, each pair shows full width at the same scale, each image links to its full-size file, and an Overlay toggle lays the built screen over the mockup at half opacity.

A branch in mockup mode that is pushed without a bound approval is held (deny) rather than advised, unless the rule is set to observe.

What approves

the person…result
picks Approve on the marked questionevery pending registration whose id the marker lists is approved
picks Request changes or Holdthe previews are dropped
does not answer (a timeout)still pending
types approve, ship, lgtm or looks good as the next prompt, after a marked question was askedapproved
types anything else next — including yes, go, ok, continuethe previews lapse

The id in the marker is what approves: it begins with the commit's sha7, which the person sees, and the registration it names is already bound to the session, the prompt and the question. The label is for the person; a question without it still approves (it did not before 2.30.0, and the person was asked twice). A marker that names no pending id approves nothing. A register that did not bind — not the last command, run in the background, with no PreToolUse stamp, or whose output does not match HEAD or this call — says so right after it runs, with the command to run and the question to ask.

Unmarked questions are ignored, so a release question answered "Approve" approves no preview, even in the same turn. Another session's answers never touch this session's previews. Registrations and approvals lapse after a day, checked whenever they are read.

AskUserQuestion accepts answers in its input, so a model could pre-answer its own question. The PreToolUse hook records, per tool_use_id, whether a marked question arrived pre-answered; such a question never marks its previews as asked (so no answer to it can approve), the PostToolUse side checks the same record again, and under deny the question is refused before it runs.

When the rule speaks

push-preview fires on a push only when all of these hold:

  • the repository has a user interface: a dev script in package.json at the root or in web/, or git config amont.agent.push-preview.ui true (false opts a repository out);
  • a pushed branch carries an interface change (below);
  • the plan the branch landed does not declare the evidence (below);
  • that commit has no approval.

Planned evidence

Some previews leave the person nothing to judge: they already decided the only visible change when they approved the plan (ADR-0028). The exemption is declared in the plan's body, so the person decides it when they approve the plan:

## Preview

evidence: the only visible change is the 28px small controls, decided in this plan.

Two hooks make it checkable:

  1. The approval record. A PostToolUse hook on ExitPlanMode reads tool_input.planFilePath when it runs (the person may have edited the plan) and, when the result has the approve shape (tool_response an object carrying plan, no is_error), writes the plan's canonical body sha (plan review: front matter and review sections left out) to ~/.claude/amont-agent/plan-approved/<sha>, journalled plan-approved.

  2. The check, at the push. For a branch whose interface change carries no picked mockup (mockup mode always asks; a screen list that cannot be taken is held as before), the evidence stands when all hold:

    • the default branch is known (checkout.defaultRemote or the single remote, then its HEAD, main or master) and shares a merge base with the pushed commit;
    • the branch adds a plan under docs/plans/ since that merge base (the first one added), and it is not a pointer (canonical:);
    • that plan, read at the commit that added it, has its canonical sha in plan-approved/;
    • its ## Preview section (outside any code fence) has a first non-empty line starting with evidence: and a reason. The template's <…> placeholder, left unedited, is not a reason.

    A pass is journalled evidence, with the plan's path, the commit and the sha. Any miss — no plan, a plan on main before the branch, a section added in a later commit, an unapproved body — falls through to the approval path, unchanged.

The record has the trust of the approvals file: a workflow aid an agent could write to, not proof. That is why every pass is journalled. The exemption lapses when the screenshots show a visible change the plan did not decide; that part is the skill's (worktree-task F7) to enforce: the hook can only see that the person approved the declaration.

What counts as an interface change

The repository-level decision above says whether a repository is gated at all. Within a gated repository, one changed file counts only when all hold:

  1. The path. It is under app/, src/ or web/ (at any depth), or it is a .tsx, .jsx, .css or .html file.
  2. The package. Its nearest package.json — walking up from the file to the repository root, read from the pushed commit (git show <commit>:<path>), not the working tree — has a dev script, or lists react, react-dom, react-native, react-strict-dom or @duro-app/ui in dependencies or as a required peer dependency. devDependencies and optional peers (peerDependenciesMeta.<name>.optional) do not count: duro-design-system's packages/cli carries @duro-app/ui exactly there, and renders nothing. A monorepo whose root has a dev script is therefore judged package by package. No package.json at all on the way up, or one that does not parse: the path decides.
  3. Not comments only. For a .ts, .tsx, .js, .jsx or .css file, the diff (git diff -U0 <base>..<commit> -- <file>; on a new branch with no base, each unpublished commit against its first parent) is read. When every added and removed line that is not blank is a comment line — starting with //, /*, *, */, or a JSX {/* … */} — the file does not count. A line with code after a closed comment (/* a */ foo()), a * line that reads like code (CSS * {), a binary file or a diff that cannot be taken all count. Any doubt counts.

git config amont.agent.push-preview.ui still overrides the repository decision; it does not change how files are judged.

A push to main or master is left to amont's branch-protect.

Supported push shapes

v1 reads an explicit remote plus explicit refspecs — git push -u origin feat/x, git push origin HEAD:feat/x, several refspecs — with -C <dir> and a leading cd honoured. A destination under refs/tags/ is excluded; one under refs/heads/ is judged, including a tag's commit pushed onto a branch.

Everything else is unresolvable and journalled with its shape: a bare git push, a remote with no refspec (git's push.default would decide), --all, --mirror, --tags, deletes, glob refspecs, a URL as the remote, a word from a substitution, an unknown flag. Under advise an unresolvable push passes with a journal note; under deny a push to a UI repository is held — --all is not the way around the gate. The shapes the journal collects decide what the next version reads.

What push-published records

It never speaks. Before a push to a UI repository the hook remembers where each branch destination stood (keyed by tool_use_id); after it, one git ls-remote per destination, and one journal line:

outcomemeaning
published-approvedthe remote now holds an approved UI commit it did not hold before
published-unapprovedthe same, without an approval — the gap advise leaves
published-no-uipublished, nothing a person could preview
already-presentthe remote held the commit before the push
unverifiedthe remote does not hold it, or could not be asked

What this is not

A registration is the agent's attestation of worktree, commit and URL, and its guide is the agent's account of what it checked. Nothing here proves the browser served that commit or that the person looked. It covers pushes made through Claude Code, not pushes typed in a shell, and it is a workflow aid, not a Git enforcement boundary.

Reading the soak

grep -E 'push-(preview|published)' ~/.claude/amont-agent/journal.log

Every line names the repository the push ran in or the registration belongs to — never the session's working directory, which may be another worktree.

The rule's own lines (registered, replaced, approved, dropped, unanswered, expired, prefilled, unbound, plan-approved, evidence) carry the answer latency on approvals and drops: option,184s,dur=0ms. The latency is the first number (184s): the hook's own clock, from the PreToolUse of the marked question to its PostToolUse. dur= is the payload's raw duration_ms, kept only for comparison — it is not the person's answer time (a minutes-long approval arrived as 0). A - in place of the seconds means the question's PreToolUse record was written by an older release. A median under about ten seconds after a week says the approval is a rubber stamp; the evidence lines say how often a plan made the question unnecessary.

The review panel

Every plan passes a panel of expert reviewers before the person is asked to approve it (ADR-0022, fleet rule work.plan-review-panel). The skill /plan-review runs the panel. The plan-review-panel rule checks, at ExitPlanMode, that it did.

Who reviews

The panel follows the areas of every repository the plan names:

areareviewers (plan-review-<role> agents)
alwayslanguage for each language, with lang=<l> in its review block, and backend
opsplatform
a command lineunix and tui
large (500 or more files or own commits)po and architect
an interfacereact, ui-design, ux-research and game-ux

Areas are read from each repository's default tree, never from keywords in the plan. The tree is the first of origin/HEAD, origin/main, origin/master and HEAD that exists, with no fetch. Each signal:

  • Languages: Cargo.toml, go.mod, package.json or tsconfig.json, pyproject.toml.
  • Interface: the ui trait in .adr.yaml areas:, or a package.json with a dev script or a runtime React or @duro-app/ui dependency.
  • Ops: a kustomization, Chart.yaml, a HelmRelease or Flux Kustomization, or *.tf.
  • Command line: the cli trait, src/main.rs, [[bin]], a package.json bin, or package main.

Vendored, lock and dist/ paths do not count. A fork's commits are counted past its upstream, and a shallow clone is judged on files alone.

What the plan carries

  • A ## Review panel section, at most 5 lines, under the H1.

  • The full reviews under a final ## Full reviews (reference).

  • A machine comment as its last line:

    <!-- panel: repos=amont-agent,decisions adds=ui reviewers=… body-sha=<12 hex> -->
    

repos= names each repository. A name resolves to git config amont.agent.plan-review.<name>.path when that is set, for a repository kept elsewhere; otherwise to $AMONT_AGENT_PLAN_ROOT/<name>, and by default to ~/Developer/Perso/<name>. adds= declares areas the plan creates: cli, ui, ops, large, or lang:<language>. A repository that does not exist yet needs a lang: entry.

Which agents to launch

amont-agent plan-panel <plan.md> answers with the hook's own code:

repos=amont-agent areas=cli,lang:rust body=cf3f43c938c4 panel=full
plan-review-backend
plan-review-language --lang rust
plan-review-tui
plan-review-unix

panel=delta lists only the reviews needed since the plan's last accepted body. panel=current lists none.

Binding a review to the plan

amont-agent plan-sha --block [--lang <l>] <plan.md> prints the review block that each reviewer's prompt carries:

<<<PLAN path=/Users/me/.claude/plans/p.md sha=<64 hex>>>>

The sha is that of what the canonical body says: the plan without its front matter, its review section, its full reviews and its machine comment, read as CommonMark. Every word, number and operator counts, code blocks and code spans count exactly as written, and so do link destinations and the kind of each block (list item, quote, heading, table cell, an ordered list's start). Blank lines, line wrapping, indentation outside code, table padding, list-marker and emphasis style and escapes do not. A landed copy in docs/plans/, which gains front matter and may pass through a formatter such as prettier, therefore hashes like the approved plan, unless the formatter rewrites code. Writing the review results into the plan never makes a review stale. Editing the body does.

The Markdown parser is pulldown-cmark, pinned exactly: a new version may read a plan differently, so a bump is its own change, and a fixture test fails when it re-hashes.

Until the transition ends (2.25 onward), a review or a baseline bound to the byte sha that 2.24 and earlier computed still counts. The hook's journal line for a pass starts match=legacy when it rested on one, and plan-sha --legacy prints that sha, for a plan whose machine comment was written before. Both go once the journal shows no match=legacy for 14 days.

A delta in a new session. A plan's baseline is found by its path, by its body, and, last, by its H1 and repos=: the same plan presented from a new session has a new file name, and after an edit only its title still names it. Two plans with the same H1 and repositories therefore share a baseline; backend still reviews the body. If the sha cannot be computed (the parser panicked), the hook asks the person instead of passing.

A review counts only when Claude Code recorded it as completed:

  • in the foreground, a result with status: completed and the agent's agentType;
  • in the background, a task notification that Claude Code queued itself.

Notification text inside a tool result, or in a message, is never read.

The declared sha stops a skipped review and a review of an older body. It does not stop a model that deliberately declares a sha for a body it did not have reviewed.

What the hook answers

situationanswer
first presentationthe whole panel, each role bound to this plan's path; backend bound to this body
body changed since the last passbackend plus every area new since then, bound to this body
same body, no new areapasses (a re-presentation, or a metadata-only write)
no machine comment, bad repos=, no review sectionrefused
no transcript, no plan path, git failedthe person is asked (ask)
refused twice alreadythe person is asked, with UNREVIEWED: refused N×
⚠ unreviewed: … in the review sectionthe person is asked

A pass is remembered under ~/.claude/amont-agent/plan-review/, which is 0700 with 0600 files, by plan path and by body. The same plan presented from a new session, under a new file name, is therefore recognised.

The default stance is deny. Set amont.agent.plan-review-panel.stance to advise to be told without being refused, or to observe to only journal.

The implementation review

Before the push of a branch that carries a plan, one independent reviewer has read the diff against that plan and the repository's active rules, and its verdict is recorded in the plan (ADR-0022, fleet rule work.implementation-review). The worktree-task skill runs the review (step F4b). The implementation-review rule checks, at git push, that it happened for the tree being pushed.

The plan step has a panel; the quality loop has one reviewer. A diff with its surrounding code is larger than a plan and arrives on every push, and the reviewer's job is narrower: findings only, never an edit, never a waiver.

What the reviewer reads

Its prompt carries, in order: the review block, Round 1. or Delta., a brief the session builds once, and its output contract. The brief is the landed plan's path and Verification section, the diff against the default branch without docs/plans/ (its stat when it is large), the repository's active constraints (aval rules --level constraint), and the plan's Non-goals. It checks plan conformance, that every recorded verification is an observable check, violated constraints (rule id and file:line), tests that cannot fail, and scope past the phases or the Non-goals.

Its result begins with the verdict, exactly, because the hook reads it:

Verdict: approve | approve-with-changes | rework
Findings:
1. [blocking|high|medium|low] <problem>. Evidence: <file:line>. Rule: <id or —>. Edit: <concrete change>.
Would still check by hand: <one to three things>

What a review binds to

The canonical tree id: the sha256 of git ls-tree -r of the commit with every entry under docs/plans/ left out, 64 hex in any object format, computed read-only. Recording the review in the plan therefore never makes it stale; a code change does.

amont-agent tree-sha [--block] [-C <dir>] [--] [<rev>]

--block prints the block the reviewer's prompt carries:

<<<TREE repo=<name> sha=<64 hex>>>>

<name> is the basename of the directory holding the repository's common .git, so a worktree checked out as amont-agent-wt-x says amont-agent. Exit 0; 1 outside a repository or for a rev that names no tree (the reason on stderr); 2 on a usage error.

What the hook does at a push

  1. Resolves the push (crate::push_target): the repository and the commit each refspec publishes. A dry run, a push to main/master, and a push of tags only are not its business.
  2. No plan on the branch: docs/plans/ unchanged between merge-base(<pushed>, <base>) and the pushed commit, where <base> is the first of <remote>/HEAD, <remote>/main, <remote>/master that exists, never the branch's own tracking ref and never HEAD. The rule declines, silently.
  3. Computes the canonical tree of the pushed commit.
  4. Looks for a binding: a completed implementation-review agent in this session's transcript whose block names this repository and tree, or a pass remembered under ~/.claude/amont-agent/implementation-review/by-tree/. Completion is read from structured fields only, as the review panel reads it; a notification echoed inside a tool result never counts. A reviewer resumed with SendMessage on a new block counts as a review of that block (below).
  5. Reads the verdict. A review that said rework and got no later delta on the same tree is not a pass.
  6. On the first transcript binding that passes, writes the pass file (0700 directory, 0600 file), so a new session pushing the same tree passes without a transcript.
situationadvise (ships)deny
reviewed this tree, verdict not reworksilentpasses
no plan on the branchsilentsilent
reviewed an older treenote, names both treesheld, names both trees
verdict rework, no deltanoteheld
never reviewednoteheld
transcript unreadable, git failed, no default branch knownnoteheld
a push shape this guard cannot read (--all, --mirror, …)journalled, passesheld

Every note ends with the next step:

amont-agent/implementation-review: no implementation review of amont-agent tree 5085edae1d66 is in this session. worktree-task F4b: `amont-agent tree-sha --block`, then launch the implementation-review agent …

A resumed review

The delta (F4b.6) may resume the same reviewer with SendMessage, giving it the new block, instead of launching a fresh one. Each round is a review of its own, read from the shape a real session recorded (tests/fixtures/sendmessage-resume.jsonl, 2026-10-08):

  1. the launch's result carries toolUseResult.agentId;
  2. a SendMessage whose input.to is that agent id, and whose result carries toolUseResult.resumedAgentId = the same id, starts a round: its prompt is input.message (the block is read from it) and its id is the SendMessage's own;
  3. the round completes only on a structured task notification whose <tool-use-id> is that SendMessage's id and whose <task-id> is the agent id. The launch's own result is never overwritten.

A message to an agent that was not launched as implementation-review, a round whose notification names another task, or a notification typed or echoed rather than queued by Claude Code does not count. Only this rule reads resumes: the plan-review panel counts fresh launches alone.

The person's overrule

A rework that survives the delta goes to the person on a marked question: its text starts with [implementation-review <repo>@<sha64>] and its options are exactly Overrule, Fix, Hold. As with the preview approval, the hook records the question before it runs and trusts the answer only when that record exists and the question did not arrive with answers pre-filled. On Overrule, and only when this session's transcript shows a completed review of that tree that said rework, the hook writes the pass file with verdict: overruled. An overrule that bound nothing says so in the model's context.

Only the hook writes a pass file

Under deny a pass file is a trust boundary. A Write, Edit or MultiEdit under ~/.claude/amont-agent/implementation-review/ is refused, with . and .. folded before the comparison; a Bash command that would write there — a redirect into it, or rm, cp, tee, python, chezmoi and the like with a word naming it — is refused too, whatever the rule's stance. Reading the record (ls, cat, find, grep) is allowed. The Bash half is best-effort by nature: a parser cannot see inside python -c beyond its words.

What the journal records

One line per push, under the rule's id, in the buckets amont-agent status counts: a declined push as unconfirmed with the reason (reviewed, overruled, no-plan, dry-run, no-branch, unresolvable), a spoken one as advised, denied or watched. The excerpt of a spoken line is structured:

tree=<12 hex> review=<stale|rework|missing|unknown> verdict=<approve|approve-with-changes|rework|->

Away from a push: overruled, answered, unbound, unknown, prefilled, unrecorded for the marked question, and denied for a refused write. The soak reads them with:

grep -E ' implementation-review ' ~/.claude/amont-agent/journal.log

Stance

Ships advise. git config --global amont.agent.implementation-review.stance deny holds a push whose review is missing, stale, rework or unknown; observe journals and says nothing. The 2026-10-06 soak review, alongside push-preview, decides between deny, advise and retiring the gate from the journal's review= counts and the first real cycles' finding quality.

The session notice

There is a mistake no command-level rule can catch, because no command is wrong: a session opens in a checkout last pulled on Tuesday, the model reads the tree it is given, and builds a feature that landed on origin/main on Wednesday. The work is correct against the code it can see.

So at SessionStart the guard does the one thing the model cannot do for itself.

The stale checkout

It refreshes origin/main — one branch, no tags, killed at five seconds, skipped when FETCH_HEAD is under ten minutes old so a burst of sessions shares one round-trip — and if HEAD is behind, says so:

amont-agent/stale-base: this checkout of amont-agent (branch main) is 8
commits behind origin/main; newest there: d3b2ed5 chore(release): 2.1.0
(3 days ago). Work that seems missing here may already exist on
origin/main — `git log HEAD..origin/main --oneline` lists it — and a
branch or worktree started from HEAD inherits the gap; one started from
origin/main does not.

It never pulls. git pull rewrites the working tree under whoever is using it, and a per-task worktree exists precisely so that nobody does that. Moving refs/remotes/origin/* is safe in every worktree at once; moving HEAD is not.

If the fetch fails or times out, the notice is computed against the last successful fetch and says so. When the checkout is up to date, or it is not a repository, or there is no remote, it says nothing.

The stale-base rule is the same fact at the moment it is about to be inherited: git worktree add, git checkout -b or git switch -c from HEAD or a local branch, while that start point is behind. The remote form (… -b feat/x origin/main) is the remedy and never fires.

git config --global amont.agent.stale-base.stance observe   # measure, say nothing
git config --global amont.agent.fetch false                 # never touch the network
git config checkout.defaultRemote forgejo                   # measure against another remote

checkout.defaultRemote is git's own key for "which remote is the remote", and a repository mid-migration — origin a mirror going stale, a second remote carrying the truth — sets it once for both git and the guard. With two remotes and no preference the guard says nothing rather than guess.

The stale guidance block

The same moment is when an agent reads AGENTS.md and believes it, and follows it for the whole session. A block generated two releases ago can be wrong before any command runs.

This one is entirely optional and entirely about amont. If — and only if — this repository carries an <!-- amont:start --> block and amont is on your PATH, the guard asks amont the question amont already answers:

amont agents-md --check

and reports drift. Two file reads decide whether to spawn anything at all, so a repository with no such block costs no process. No amont on PATH means this says nothing — which is the right answer for anyone who does not use it.

It reads amont's stderr, not its exit code, and that is not fussiness: agents-md --check exits 1 for a file it could not read as well as for a drifted one, so the exit code alone would announce staleness for a permissions error. A session-opening notice that cries wolf is worse than one that occasionally says nothing.

git config --global amont.agent.agentsMdNotice false   # silence this half

Open plans

If the session opens inside a repository with a docs/plans/ directory (ADR-0022), one line per active plan: its title and its first unchecked phase (- [ ] …). A pointer file — a plan whose canonical copy lives in another repository — prints the canonical path and the phases this repository carries. At most five lines; the rest are counted.

amont-agent/plans: 1 open plan in docs/plans/ (read one on demand; they are history, not instructions):
- Settings toggle (2026-09-28-settings.md) — next: Phase 2 — route

Closed plans (done, abandoned, imported) print nothing. A plan with no status, or one outside that list, is reported: nothing else reads plan front matter, aval check included. Nothing from a plan's body beyond those two lines is loaded — an old plan is history, and loading it into every session would compete with the instructions that are current.

Configuration

Every key is plain git config, readable and removable without this tool.

The kill switches

keydefaulteffect
$AMONT_AGENT_OFFunsetany value switches the guard off for that shell
amont.agent.enabledtrueswitches it off everywhere

The environment variable is checked first, before any rule runs, because reading a git config key costs a process and the variable costs nothing. amont.agent.enabled is consulted only once something has already fired.

Both take git's own boolean dialect — true/false, yes/no, on/off, 1/0, case-insensitively — because git parses them, not us. A value git refuses warns once and falls back to the default rather than failing.

Stances

keytakes
amont.agent.stanceobserve | advise | deny — every rule
amont.agent.<rule>.stancethe same, for one rule

Most specific wins. See stances.

The session notice

keydefaulteffect
amont.agent.fetchtruemay a session opening touch the network at all
amont.agent.agentsMdNoticetruemay it mention a stale amont guidance block
checkout.defaultRemote—git's own key; which remote is the remote

What is deliberately not configurable

Which file any of these keys can be set in. --global, then --system, and nothing else. Not a committed file, and not the --local/--worktree config of the repository you happen to be standing in — see stances. Write them with git config --global; a bare git config inside a repository writes --local, where this tool will not look.

Whether a failure is silent. It always is. See what it will not do.

Whether the journal can affect a decision. It cannot.

What it will not do

It never emits allow. That would short-circuit your own permission prompt, so a guard approving everything it has no objection to would have switched off the permission system it was installed beside. Silence is how it says "no objection".

Every failure path is silence. An unreadable payload, an unknown event, a command it cannot parse, a rule that panics, a journal it cannot write — all of them exit 0 having written nothing.

A hook that fails toward refusing gets in the way of work you knew was correct, and the fix people reach for at that moment is to delete it from settings.json, which switches off every rule at once. One that fails toward silence loses a single firing. That trade is the whole posture, and it is why the hook payload is parsed with serde_json::Value and hand-written accessors rather than a derived struct: a field that is missing or has changed type becomes "no opinion", not a parse error somebody would be tempted to treat as an opinion. What the hook writes back is emitted by a hand-rolled escaper (src/json.rs), so the reading and the writing share no representation.

It does not judge what it cannot read. Heredocs without terminators, unbalanced quotes, eval — all opaque, and opaque never fires.

The unit of that is the pipeline, not the line. | chains one command's output into the next, so a stage we cannot read makes the whole run unreadable. &&, || and ; do not: a command there is as independent of an unreadable neighbour as of any other clause, and treating the whole line as opaque cost every rule on it — measured over 33,774 real commands, 218 had a readable pipeline thrown away. Those are judged now, and check names the run it could not read beneath the verdict.

Two things keep that honest. eval, source and . run in THIS shell and can move it, so one of them still hides the whole line — otherwise a later confirm would resolve a path against a directory we can no longer vouch for. And a finding from a partly-read command never refuses: it advises, whatever stance the rule carries. Total opacity would have let the command run, so blocking on half a reading is the worst outcome available.

It does not phone home. No telemetry, no update checks, no fetches — with one exception, which is git fetch against your own remote for the session notice, and amont.agent.fetch false switches that off.

No repository can change a stance. Stances are read from --global and --system git config only — never from a committed file, and never from the .git/config of the repository the agent is standing in, which is a file that agent could write. See stances.

The journal

Every firing is recorded at ~/.claude/amont-agent/journal.log, redacted, and never transmitted anywhere.

It only counts. Nothing in it may participate in a decision — the rules read the command in front of them and nothing else. A guard whose verdict depended on its own history would be one you could not reason about from the command alone, and could not test from a corpus.

Shell analysis

src/analysis/ answers one question about a single Bash command, before it runs: which network transfers it may make, and which it must, with the evidence for each count. request-fanout is its first consumer. The analysis decides nothing; the rules do.

Layers

layermodulejob
frontendanalysis::frontendsource → crate-owned IR, spans on every node
semanticsanalysis::interp, state, domainabstract interpretation: state, control flow, exit statuses, counts
command modelsanalysis::modelsa client's arguments → its explicit transfers
policyrules::request_fanoutwhether the result warrants advice, and the words

Everything in analysis is pure. The dialect is an input; nothing reads the environment, the filesystem or the network. The legacy lexer (src/shell.rs) is untouched and still serves every other rule.

Inputs

  • The command text.
  • The dialect — bash, zsh or unknown. The hook reads it from SHELL; check --dialect sets it; the backtester always uses unknown, because transcripts do not record the shell. Where bash and zsh differ, an unknown dialect gets the join of both readings — never one of them silently.

Dialect differences that are modelled: whether the last element of a pipeline runs in the current shell (zsh yes; bash only under shopt -s lastpipe); word-splitting of unquoted parameters (bash splits, zsh does not); an unmatched glob in a URL-shaped word (bash passes it through; zsh refuses the command); break/continue with no loop in the current shell, as inside ( … ) (bash warns and carries on; zsh leaves).

Where the command does not say which seq runs, seq 5 1 is either nothing (GNU) or five numbers (BSD, macOS): both are allowed.

tools/shell-oracle/ checks all of this against bash and zsh actually running generated scripts with stubbed clients — see its Cargo.toml.

Assumptions

Every "established" count holds under these, and the rule's message lists them:

  • the shell is not killed from outside;
  • redirections succeed;
  • errexit and pipefail are off unless the command sets them (then they are modelled);
  • no aliases are defined;
  • inherited environment variables are unknown;
  • the command has no positional arguments of its own.

What is counted

Explicit command-line transfers: the requests a client's arguments ask for. Redirects, authentication hops, a client's own retries and pagination, and defaults from ~/.curlrc are named as possible extras and never counted.

A count is an interval: at least lower, at most upper, where upper is a number, saturated (too large to represent), uncapped (positive evidence that nothing bounds it: a loop following next-page links, --paginate, wget -r, while true with no way out), or unknown (no evidence either way). Zero annihilates: code that provably never runs makes no transfers, whatever it would have done.

A finite upper bound on a loop comes only from the counter pattern: a literal initial value, a comparison as the condition's final command, exactly one unconditional +1 of the counter in the body, and no other write to it — including from a function the body calls, or from anything unmodelled — and no continue that could skip the increment.

Pacing

Each path through the interpreter carries how long it definitely slept (sleep 30, sleep 2m; a backgrounded sleep does not count). A loop is paced when every path back to its head — the end of the body, or a continue — slept: a continue that skips the sleep, or a sleep behind a condition, leaves it unpaced. A paced call site records the interval and its burst: the transfers it makes per iteration, an inner loop or --retry included. The innermost paced loop is the one recorded.

Paced and unpaced transfers never join across if branches: a poll in one branch and a one-off request in the other are two different kinds of load.

request-fanout exempts a call site paced at least 10 seconds apart whose burst is within its budget: a poll, not a burst, and foreground-poll's business if it runs in the foreground.

The supported subset

Sequences, &&/||, pipelines and !, if/elif/else, case with ;;, ;& and ;;&, for, for ((…)), while, until, { }, ( ), (( )), [[ ]], [ ]/test, function definitions in both forms, local/typeset, break/continue with levels, return, exit, set -e, set -o pipefail, shopt -s lastpipe; $( ) and backticks, $v, ${v}, positional parameters, $@/$*, the ${v:-x} family (not evaluated: its value is unknown), every quoting form, unquoted brace expansion.

Unknown and Incomplete

  • Unsupported constructs — eval, source, trap, xargs, parallel, a command name built at run time, a program that runs another program — may do anything: every variable becomes unknown, the command may exit or never end, and the region is reported as activity the analysis could not follow. It never erases a known count elsewhere in the command.
  • Incomplete means a resource limit ended the analysis (nesting depth, IR size, interpreter steps, call depth, recursion). What was found before it is kept; nothing after it is known.
  • Cardinalities — {1..999999999}, seq, curl URL ranges — are computed arithmetically and never expanded.

A rule built on the analysis never fires on an unknown count alone.