Environments & Deploy
Current state (three environments)
- prod —
mainbranch → Vercel projecttryvio-shopify→app.tryvio.ai(prod Supabase, prod Shopify appe0ee…). - dev —
devbranch → Vercel projecttryvio-dev→dev.app.tryvio.ai(dev Supabasexcyrjksywrdmerbocwjf, dev Shopify app3f90…,SHOPIFY_BILLING_TEST=true, and mock try-ons by default viaTRYON_MOCK_MODE—env-driftnow enforces that switch’s presence on dev after every dev deploy (#568)). - staging — not live yet; a prod-copy + PII-scrubbed rehearsal environment inside the release train, in progress as epic #376.
Deploying the web app
Merging to main auto-deploys prod, since 2026-07-14: GitHub Actions’ production.yml runs
verify (tsc + full vitest) → deploy to Vercel → post-deploy health / schema-drift / env-drift /
prod-smoke. Manual deploy is only the fallback for when Actions itself can’t run — e.g. the
Actions minutes quota is exhausted (the verify job goes red with 0 steps, not a code failure;
this happened on the 2026.35 ship, 2026-08-20).
# preferred fallback — runs the SAME verify gate locally, then builds remotely on Vercel
node scripts/deploy-local.mjs --target=prod --confirm=SHIP-PROD
# simplest raw fallback (from apps/web, no gate)
cd apps/web && npx vercel --prod --yesThe deploy aliases to app.tryvio.ai. The .vercel/ link in apps/web targets the
tryvio-shopify Vercel project.
Deploying the docs sites
# public customer docs → docs.tryvio.ai (Vercel project "docs")
cd apps/docs && npx vercel --prod --scope mihogv-1894s-projects
# internal docs (gated) → Vercel project "docs-internal"
cd apps/docs-internal && npx vercel --prod --scope mihogv-1894s-projectsdocs.tryvio.ai needs a DNS record at Namecheap (A docs 76.76.21.21). The internal docs
stay behind Vercel Authentication (no public domain).
Deploying Shopify config
When shopify-app/shopify.app.toml changes (scopes, webhooks, app proxy):
cd shopify-app
npx shopify app deploy --allow-updatesAdding a new scope emails installed merchants to re-approve permissions; they keep old permissions until they accept.
Deploying the storefront widget (manual — and only from the target branch)
The theme app extension is not in the pipeline:
cd 02_app/shopify-app
node scripts/deploy-widget.mjs dev # → DEV Shopify app (dev URLs swapped in transiently)
node scripts/deploy-widget.mjs # → PROD Shopify app (needs the owner's шипвай)The tool publishes whatever is in the current working tree — that is the whole hazard. Under
Process v4 every work branch is cut from origin/main, whose widget assets lag dev by a full
release, so a deploy from a work folder used to silently roll the widget back and report success.
It happened on 2026-08-22: the dev app served a widget without #410’s fix for ~20 minutes, caught by
a human noticing, not by a gate. Because a widget version can stay pinned on a storefront for
weeks, the same mistake on prod is worse than a bad deploy.
Since #434 scripts/widget-deploy-preflight.mjs runs first — before the lock, the rebuild and the
minify — comparing the tree against the branch the target app must serve (origin/dev for dev,
origin/main for prod) over the paths the deploy actually publishes (both extensions plus the mode’s
shopify.app*.toml). It prints branch, commit, target and a hash per published asset, then:
| tree vs target | DEV app | PROD app |
|---|---|---|
| at target (byte-identical) | deploy | deploy |
| behind (target has widget commits the tree lacks) | refuse — names the files + missing commits; --allow-rollback --confirm=<head sha> publishes it on purpose | refuse, no override |
| ahead (unlanded widget work) | warn, deploy | refuse |
| uncommitted published files | warn, deploy | refuse |
target ref unresolvable / git fetch failed | warn (refuse if unresolvable) | refuse |
--preflight-only runs the check and publishes nothing. A difference confined to docs (*.md) or to
the generated assets/tryvio-modal.js (rebuilt from src/ on every deploy) never refuses. Regression
harness: node scripts/widget-deploy-selftest.mjs — 35 checks in a throwaway repo, including an
end-to-end run with the Shopify CLI stubbed and a check that pins the pre-#434 behaviour.
Branching model (Process v4, owner decision 2026-08-20 on #387, built by #396/#397)
The release is assembled by merge all week, not by cherry-pick on Monday (that was Process v3, below).
Branch types. feature/<issue>-<slug>, bugfix/<issue>-<slug> (a bug
riding a release), hotfix/<issue>-<slug> (P0, straight to main), chore/<issue>-<slug>,
docs/<issue>-<slug>, and release/<year>.<week>. The issue number is
mandatory in every work-branch name. node scripts/agents/worktree-new.mjs <issue> creates
the worktree + branch — type comes from the issue’s labels, slug from its title, rerere
enabled. branch-guard denies creating a work branch without an issue number.
Claiming: the state guard and linked packages (#438)
worktree-new.mjs refuses work that already exists. An issue labelled on-dev or
awaiting-prod-approval, or one that is CLOSED, is not claimable: the tool aborts before the
worktree and before any board write, prints the issue’s state, the last 🌿 claim comments, and what
to do instead (QA it on dev; a real defect in it is a NEW bug issue on a bugfix/ branch). Until
this existed the tool printed a warning for the closed case and nothing at all for the landed one
— on 2026-08-23 a console re-opened a worktree for #421, which had been finished, QA’d 6/6 and
labelled awaiting-prod-approval the day before. --force-claim (alias --force) overrides every
claim guard for a deliberate repair, and what it overrode is written into the 🌿 comment — an
override nobody can see is the same as no guard.
A link between two issues is a FIELD, not prose. Ships with: #N on its own line in the issue
body is the one machine-readable statement that two issues are one change split in two. One parser
now reads it everywhere — scripts/agents/_issue-links.mjs, used by worktree-new.mjs,
land.mjs, board-sync.mjs and the release manifest:
- it is line-anchored (bullets, block quotes, brackets and markdown emphasis are fine). A
Ships with:quoted mid-sentence or pasted inside a code fence is prose and creates no dependency — #438’s and #450’s own bodies used to giveland.mjsphantom deps this way; - it reads the whole id list:
Ships with: #420, #406used to be read as#420alone, leaving the other half of the coupling silently unenforced.
Claim the package, not a half. worktree-new.mjs <A> announces the package, its ship
order and the exact command to take both in ONE worktree:
node scripts/agents/worktree-new.mjs <A> --with <B>. That claims both issues, points
both 🌿 comments at the same folder and names the branch after the issue that must land first
(land.mjs refuses a follower until its declared half is on the target). A mutual declaration
has no ship order, so the tool says so at claim time and tells you to keep the declaration on the
follower only — a package worked in ONE branch and promoted by ONE merge is still the simplest
shape. It is no longer a deadlock, though: see A mutual declaration is a peer link below.
The board shows a package as one row. node scripts/agents/board-sync.mjs --ready
(--milestone S<year>.<week> for the sprint) collapses a linked package into a single
row with both numbers, the pooled priority/zones, the ship order and the --with command;
gh issue list --label ready cannot, which is how #420 and #421 were handed to two consoles.
The nightly board-sync.mjs --lint-links labels every package member linked, comments on the
half of a one-sided link that is silent (with the exact line to add) and reports a Ships with:
written in prose. Regression harness: node scripts/agents/linked-issues-selftest.mjs (hermetic —
throwaway repo, stubbed gh via TRYVIO_GH_STUB).
A mutual declaration is a peer link, not an order (#450). Until 2026-08-29, Ships with:
was read purely as a sequence: land #N first. Two issues declaring each other therefore had no
satisfiable order at all — land.mjs 441 said “land #439 first”, land.mjs 439 said “land #441
first”, forever, and each of two consoles concluded the other issue was the blocker while nothing
went red. It hit #439/#441, a money change, which is exactly the shape where “must not ship apart”
is the point. land.mjs now reads the partner’s OWN body, confirms the back-reference, and treats
the pair as peers:
- on
deveither half may land first (nothing ships fromdev), and the run says#A and #B declare each other — a Ships-with CYCLEinstead of pointing at the other issue; - promoting into the release is allowed while the other half is coming — already on the train
or labelled
awaiting-prod-approval— and prints what is still owed and by whom. If it is neither, the promotion ABORTS naming the cycle, never “promote the other half first”; - the freeze is where “ship together” is enforced.
release-cut.mjs --freezeis a hard RED when the train carries one half of a declared pair, and it is the one layer that cannot deadlock: by freeze time every issue is provably either on the train or off it. Fix it by promoting the missing half, or drop this one withrelease-open.mjs --recompute --exclude <issue>.
A one-directional dependency is unchanged — still strictly ordered — and so is an issue that
merely follows a cycle without being in it. Regression harness: node scripts/agents/land-selftest.mjs
(T24–T28).
Merge direction. Work branches are cut from main and merge only main (to catch up
after a hotfix) — never dev, never release/*, never another work branch. branch-guard
denies those merges: a dev merge into a work branch would ride the release the hour that branch
is promoted.
Landing on dev. node scripts/agents/land.mjs <issue> merges the work branch into
dev inside a temporary detached worktree: rebuilds the widget bundle if its source changed,
runs proof-lint + real-seam (for money/DB issues) + pnpm/tsc/vitest on the merged tree,
publishes dev, and labels the issue on-dev. Conflicts are resolved in that temp worktree and
recorded by git rerere — the work branch itself is never modified. A new conflict is a
CONFLICT-STOP that keeps the worktree for you to resolve, commit, and re-run with --continue.
Resolve ONLY the conflicted files there (#508): the merge’s tree is compared with the mechanical merge
of its two parents (git merge-tree), and any other path that differs — a follow-up fix typed into the
merge — is REFUSED, because a promotion merges the branch and would never carry it, and a re-merge would
drop it. Put such a change on the branch. --merge-only-ok "<why>" keeps a deliberate one: it is
recorded as a Merge-only-edits: trailer, replayed onto every re-merge, and named again when the issue
is promoted. node scripts/agents/merge-only-edits.mjs [--issue N] lists the ones already on dev
(8 of the last 150 land merges on 2026-09-30, #705’s stylesheet lines among them).
If dev moved underneath (another console landed first), the tool re-merges and retries — see
below. dev is merge-only and derived — no direct commits.
A rejected push says WHICH kind of rejection (#485)
A busy dev used to be able to starve a land: the retry threw the temp worktree away, so every
attempt paid another pnpm install + tsc + full vitest (~7 min), and three of those was the end of
it. Worse, every non-zero git push was read as the race — including a push refused by the
pre-push landing gate (#398), which was then reported as “dev moved while we were working” and
re-verified three times while dev had provably not moved. Nine attempts, ~60 minutes of green
verify discarded, and the actionable two-line fix scrolled off the screen.
scripts/agents/_push-classify.mjs now decides, from git’s own ref-status wording:
| kind | git said | what happens |
|---|---|---|
RACE | (non-fast-forward), (fetch first), (stale info), (incorrect old value provided), cannot lock ref … but expected | re-merge onto the new head and retry |
BLOCKED | anything else — a pre-push gate red, pre-receive hook declined, auth, network, a protected branch | STOP after ONE attempt, name the cause, re-print git’s own output at the bottom, state that the branch did not move |
It is the mirror image of #435’s verify classifier on purpose: there an unrecognised red falls back
to ASSERTION because retrying forever is the worse harm; here the retry is the harm, so a
retry needs positive evidence and everything unrecognised is BLOCKED.
A RACE retry is now cheap and honest:
- the temp worktree is re-merged in place, keeping
node_modulesand the.land-verified-<tree>certificate — so a re-run after an abort resumes instead of starting from zero; - if the commits that beat us touch no file in this land’s own diff, the merge is republished
without re-running the verify and without #414’s
Verified-on-merged-tree:stamp, sodevelopment.ymlverifies the combination in CI. The certificate is never claimed to cover a tree it does not (#431 is untouched:verifiedTreeisnullon that path); - if they do overlap, the full verify runs again — correctness is not traded away;
- publishing is serialised by a short
land-publishlock (lock.mjs) taken immediately before the push and released immediately after — never across the verify. It is fail-open: a lock that cannot be taken withinTRYVIO_LAND_LOCK_WAIT_MS(120 s) is a WARN and the push happens anyway, and a holder older than 10 min is broken as stale; MAX_ATTEMPTS(3) bounds the expensive attempts;TRYVIO_LAND_MAX_ATTEMPTS(8) bounds the cheap ones. Giving up prints how long it spent, one line per attempt naming what beat it, and the resume command.
Regression harness: node scripts/agents/land-selftest.mjs (T7, T16–T20).
A promotion releases the worktree’s dependencies (#448). The lifecycle (owner, 2026-08-28) is:
the console finishes → the issue is labelled awaiting-prod-approval → its node_modules are
purged, and the tree stays → at the SHIP the worktree itself goes. land.mjs <issue> --into release performs the purge on a successful promotion — that command already refuses to run without
the label, so the trigger is the lifecycle. It never fires on a dev land (QA on dev still wants an
installed tree) or on a CONFLICT-STOP/abort (that tree is where you work next); --keep-modules
opts out. Dep roots are discovered, not listed — the primary checkout carries eleven of them —
and deleted with nativeRm, so the pnpm JUNCTIONS into the shared store are removed as links.
Following one wipes that store for every project on the machine.
Already-parked deps: every worktree-gc.mjs run now reports them (node_modules: N dep dir(s) parked in M of the listed worktrees), and worktree-gc.mjs --modules [--yes] — or modules-gc.mjs
standalone — reclaims them without removing any worktree, branch or registration. It defaults to
refusing: dirty, git- or agent-locked, process-held, active in the last 2h, an issue not yet
awaiting-prod-approval/closed, or simply unknown (gh unreachable) all keep the tree, named. Before
this, worktree-gc printed “42 listed, 1 eligible, est. ~0 MB reclaimable” on a 98%-full disk
while 37 dep dirs sat inside trees it was correctly protecting.
🚨 Never size a pnpm tree by summing file lengths. pnpm hardlinks into a shared store, so a per-file walk counts the same blocks once per worktree. Measured 2026-08-28: a fully installed worktree is 387 MB apparent and ~9 MB of actual disk (sampled files carried 16-22 hardlinks). Both figures this issue was justified with (10.9 GB, then 7.0 GB) were that over-count — purging worktree deps does not fix a full disk. Sizes are reported as the volume’s free-space delta.
A red verify says WHICH kind of red (#435)
land.mjs and deploy-local.mjs stream and capture their gate output, then classify a failure
before aborting. The two kinds demand opposite responses, so they never share a sentence or an exit
code:
| kind | means | abort line | exit | retried? |
|---|---|---|---|---|
ASSERTION | the tree is broken — a type error, a failing test, a bad lockfile | ... ASSERTION-RED — the CODE failed | 1 | never |
RESOURCE | the machine ran out — OOM, ENOSPC, EMFILE, SIGABRT, a vanished temp tree, a vitest timeout with no failed assertion | ... RESOURCE-RED — the MACHINE failed, not the tree | 75 (EX_TEMPFAIL) | yes, bounded |
The signatures live in one shared helper, scripts/agents/_verify-classify.mjs, which both tools
import — never copy-pasted patterns. Precedence is one-directional: a genuine failing assertion
always wins, so a real red produced while the machine also happens to be starving is still
ASSERTION. An unrecognised failure falls back to ASSERTION (fail-closed) and is never retried.
Both aborts print the machine state at the moment of failure (free RAM, free disk, node process
count), and land.mjs warns up front when free RAM < 1 GB or free disk < 5 GB — thresholds
overridable with TRYVIO_VERIFY_MIN_RAM_MB / TRYVIO_VERIFY_MIN_DISK_GB. A RESOURCE red is
retried once after a 90 s pause (--resource-retries <n>,
TRYVIO_VERIFY_RESOURCE_PAUSE_S); 0 disables it. Regression harness:
node scripts/agents/verify-classify-selftest.mjs.
…and a GREEN verify must prove WHAT it measured (#436)
status === 0 was itself a signature of nothing. A vitest run whose workers Windows refused to spawn
exits 0 with Errors 124 and Test Files 245 passed | 2 skipped (248) — 103 of 351 collected
files never ran, and the summary still ends in passed. Both tools therefore call gateVerdict()
instead of reading the exit code: it answers “green” only for a step that can show what it covered —
no runner Errors, the summary’s buckets adding up to the number collected, and (for the test step)
a summary printed at all. A gap is a RESOURCE red like any other, so land.mjs writes no
.land-verified-<tree> certificate for it: one degraded measurement never becomes durable evidence.
Two more things changed with it:
TS6053is no longer decided by the filesystem alone. On 2026-08-27 a second process removed a land worktree mid-compile;tsclost its ownlib.dom.d.tsand every file the include pattern matched, and the classifier — checkingexistsSyncafter the tree had been re-created — answeredASSERTION-RED — the CODE failed, the one verdict that is never retried, over code that was fine. Two text-only discriminators now pre-empt that: a missing compiler lib (nothing in this repo can cause it, and every other diagnostic in such a run is an artefact of it), and an output whose diagnostics are allTS6053over five or more files — a glob-expandedincludecannot produce that, so nothing was compiled at all. One or two genuinely missing files stay a code red.- Gate children run under an explicit heap ceiling —
--max-old-space-size, 4096 MB by default,TRYVIO_VERIFY_HEAP_MBto change it — instead of whatever V8 sizes itself to from the machine’s total memory. A ceiling the caller already set is never overridden.
Promoting to the release train. QA’d on dev → label awaiting-prod-approval →
node scripts/agents/land.mjs <issue> --into release merges the branch into the currently
open release/<year>.<week>. Checks: the label is set, the issue has no unpromoted
dev-only lineage, and any declared Ships with: #X dependency is already on the release (unless
#X declares us back — see below); rerere replays the conflict resolution already made when the
branch landed on dev. The release branch
itself is opened at main right after the previous ship, by
node scripts/agents/release-open.mjs; release-open.mjs --recompute --exclude <issue>
rebuilds it as main + the remaining promotions — scope removal without reverts,
force-with-lease — which voids, by design, any approval nonce bound to the old head.
A package promotion (land(620,637): …) is ONE merge, so --exclude of either issue drops the
whole merge and the tool names the partner that goes with it. Each rebuilt subject is re-derived
from the commit scopes the promoted branch declares, never replayed verbatim — a promotion whose
subject under-named an issue (land(512) carrying fix(513)) comes back as land(512,513) (#530).
The release board (#739). Every release has ONE private Claude artifact the owner reads instead
of the board: the issues on the train (each with a one-line Bulgarian summary, risk chips, how and
where it was tested and up to 10 of its proof screenshots), what is on dev waiting for his OK, what
he approved that nobody promoted, what is in progress, outreach (label outreach — it runs on the VM
from dev and never rides the train) and issues whose work already shipped but are still open. Lanes
come from facts, not labels: promotions on release/<w>, dev lands, earlier ships on main — the
first build found 36 of 51 “approved” issues already on prod. The train position (collecting → frozen →
staging → waiting for шипвай → on prod → closed) comes from the release PR, the staging record and the
tag, and a tag alone never reads “closed” (a selftest once pushed v2026.40 onto the 2026.36 merge).
node scripts/agents/release-board.mjs builds it into .release-board/<w>/ (gitignored; screenshots
converted once per blob into ~/.tryvio/release-board/cache) and prints the exact Artifact publish
step — only a Claude console can publish an artifact, so a PostToolUse hook
(.claude/hooks/release-board-hook.mjs) tells the console that just ran a land, a promotion, a
flow-label edit, the freeze, approve, ship, prod-verify or a close to rebuild and republish. The link
lives in ~/.tryvio/release-board/boards.json (--set-url after the first publish, --published after
each one), and playwright-debug/notify-telegram.js appends 📋 Релийз <w>: <link> to every
message and caption it sends (once; TRYVIO_TG_NO_BOARD=1 opts out). Bulgarian summaries:
release-board.mjs --summary <issue> "<text>" — the build lists the ones missing.
Monday freeze. node scripts/agents/release-cut.mjs --freeze --week <year>.<week>
does no assembly (the release is already built by the week’s promotions). It derives the shipped
issue set from the land(N): merges, runs content-sanity vs dev (dev may be ahead on a file
only because of unpromoted work — a difference with no such dev commit is a STOP), the widget
bundle check, the Ships with: package gate (a HALF package on the train is a hard RED, #450),
proof-lint --prod-verify-strict, real-seam, tsc + vitest, writes the manifest, and
prints the gh pr create --base main --head release/<w> step.
The manifest is COMMITTED (#482). RELEASE-<w>.md is the machine-written record of what
the owner approved — the promoted issues with their proof links and zones, the gate verdicts, the
migrations and env keys in scope — so the freeze commits it onto the release branch and it rides
the release PR to main, arriving with the thing it approved. It sits at the repo root, which the
landing gate’s GATED_RE does not cover (a manifest-only push is issue-refs: N/A), and it is a
plain commit, which pr-purity ignores on a release branch (that job only inspects merges). The
content gate skips it by name — it is release-only by construction, so without that a
--freeze --resume would abort on the manifest the first freeze committed. Re-running the freeze
re-writes it and commits only if it changed. Before this, the file was written into the temp
worktree that release-close.mjs unregisters and never staged: after the 2026.36 ship the only
copy on earth sat in an orphan directory that worktree-gc was holding by age alone.
release-close.mjs now checks that the manifest exists on origin/main or origin/dev (root or
evidence pack) before it removes the worktree; if not, it rescues the file into
05_tasks/proof/release-<w>/ on dev, and if even that fails it keeps the worktree rather
than deleting the only copy.
Shipping. The release PR merges to main only on the owner’s explicit “шипвай”
(approve.mjs nonce — golden rule #1, no exceptions). Hotfix lane (P0 only):
hotfix/<issue>-<slug> → pr-prod.mjs → шипвай →
node scripts/agents/backmerge.mjs, which merges main into both the open release and dev in
one command. After the ship, node scripts/agents/release-close.mjs --week <year>.<week>
tags v<year>.<week> on the PR merge commit, back-merges main → dev, deletes
release/<w> and its worktree, deletes every origin work branch already merged into main,
and opens the next release — derived from the week being closed (2026.36 → 2026.37, year
rollover included, _week.mjs). It used to call release-open.mjs with no week, whose default is
“the ISO week of next Wednesday”; on 2026-08-26 that resolved to 2026.35, hit a leftover branch, and
the next release was simply never opened (#455). An already-open next release is now reported as
such instead of as a generic failure.
The chain checks itself before it moves anything (#455). ship.mjs lints its own modules —
_ship-steps, _manifest, _staging, prod-state, ship, release-close, prod-verify,
deploy-local — for names used in call position that resolve to nothing
(_step-refs.mjs), and aborts naming file, line and identifier. This runs before step one, dry
run or not. It exists because scripts/agents/** is outside the web tsconfig and has no vitest
coverage: on 2026-08-26 stepDeploy called an undefined gitOut, --dry-run reported
“12 steps green” that morning (a dry run executes nothing), and the first evaluation of that line
happened one step after the irreversible merge — prod served the old code against the new schema
until a human ran deploy-local.mjs by hand. stepDeploy also fetches before pinning the Actions
run to origin/main: the merge happens on the remote, so a stale local ref points at the previous
release’s commit, whose production run is green and would read as this ship’s deploy.
CI enforcement. On a release/* PR, pr-purity accepts only land(N) promotion merges (of
pure work branches) plus main back-merges. The >30-commit fence applies to hotfix PRs only.
The proof-lint job runs for release/* and *-prod heads. A red pr-purity/proof-lint check
BLOCKS the merge (#398 AC3 — “block, not delete”): branch-guard reads gh pr checks before it
honours the owner’s nonce and denies on a failed check; a check that has not run (Actions minutes
exhausted, pending) is printed as a visible WARN: pr-purity did not run — never a pass.
ship.mjs reads the same two checks in its own merge step (#753): a merge spawned from a script
never passes the hook, and release 2026.39 went to main with purity red. A red check stops the
chain there and the owner’s approval stays unburned.
The release’s own promotions are not lineage (#753). Work is cut from the open release, so a
promoted branch carries that release’s earlier land(N) merges. pr-purity skips an inner merge
that is already in the release as it stood before the promotion (<promotion>^1); anything else
whose second parent is not on main is still an error. The rule is foreignMerges in
scripts/agents/_lineage.mjs, which land.mjs calls; the workflow is bash, so
pr-purity-selftest.mjs runs the workflow’s own script and that function on the same commits and
reds when they name different merges.
Landing gate — gates run where the work lands (#398)
Until 2026-08-20, 25 of 27 gate points ran Monday-or-later and every freeze opened with ~20
red-on-form results (stamps, waivers, a filename range read as a phantom file, five test-only env
keys) discovered under time pressure in a session that had not written them. Now one offline
bundle, scripts/agents/landing-gate.mjs, runs on every push:
- worktree pre-push hook — an ABSOLUTE wrapper at
.git/ttn-gates/hooks/pre-push, installed byworktree-new.mjsand repaired bygates-doctor.mjs --apply(#430; the old committedscripts/githooks/pre-push+ relativecore.hooksPathis still the fallback in a checkout that has one). A red blocks the push and names the exact line to fix.git push --no-verifyonly moves the red to CI. - CI —
.github/workflows/landing-gate.ymlon every push todevandrelease/**(no path filter; one job, noapps/webinstall, ~1 min). Writes a PASS/RED/WARN table to the step summary AND the job log, and one::error::annotation per red gate carrying its fix line. The step runs underbash -e, so the gate’s exit code is captured rather than allowed to end the step — until #629 a red ended it before any of that printed, and the only line in the log wasProcess completed with exit code 1. A gate that crashes without JSON still keeps its exit code and says so.
| gate | what it checks | red = |
|---|---|---|
branch-name | v4 work branch is <type>/<issue>-<slug> | rename command printed; legacy feat/x → WARN |
issue-refs | the push DECLARES its issue(s): scope fix(358):, #358 in the subject, Refs: #358 trailer, a land(358): merge, or the branch name — only when gated paths change (02_app/, scripts/agents/, .github/, 05_tasks/proof/) | add Refs: #<issue>; docs-only pushes are N/A; Merge pull request #N is NOT a ref |
proof-lint | proof-lint over the push’s issues + every proof folder the push touches; work branch: an existing proof must be clean (missing = reported, land.mjs requires it); dev/release: it must exist; release: --prod-verify-strict | every line ends with → add to PROOF.md: "<literal>" |
env-classify | every env key the diff ADDS a read of is in env-registry.ts — or is test-only by verified reads (_env-scan.mjs, #388: every read in *.test.ts(x) / __tests__/); a registry test-only claim contradicted by a shipping read is red; since #443 a required/waived key with no ENV_ORIGIN entry (how to obtain its value) is also red. Since #569 the gate’s own vocabulary: the pre-push gate runs from the origin/dev snapshot, so a level it does not know is RED UNKNOWN LEVEL — except a level the push itself adds to BOTH REGISTRY_LEVELS (_env-scan.mjs) and EnvLevel (env-registry.ts), which is accepted as classified and reported as a WARN VOCABULARY "<level>" (its meaning is judged by the pushed tree’s own tools, e.g. landing-gate.yml) | the registry line to paste is printed |
real-seam | real-seam-gate.mjs existence check for money/DB diffs (#263) | rs-test or Real seam: N/A — … |
widget-bundle | a widget-src diff rebuilds tryvio-modal.js + check-modules; the committed bundle must match | commit the regenerated asset; no esbuild → WARN did not run |
selftests (#446) | a dev/release/* push touching scripts/agents/** runs the agent tooling’s own proof harnesses — the set is the REGISTER scripts/agents/_selftests.mjs, linted against the directory. N/A on a work branch (its main-based copies are LAST release’s, so their red is a fact about the branch, #430 — land.mjs runs the check on the MERGED tree). Two members are excluded in writing: selftest-worktree-new-claim.mjs writes REAL GitHub issues, deploy-local-selftest.mjs dies on the local Node runtime (#433/#435) | the harness’s own FAIL lines + reproduce: node scripts/agents/<harness>; an UNDECLARED harness is red with the register line to paste; a hung harness is killed and red; --no-selftests (env spelling TRYVIO_LANDING_GATE_SELFTESTS=skip, the one a pre-push hook can reach) is a WARN, never a pass |
Setting a missing key is no longer a human opening the Vercel dashboard: node scripts/agents/env-set.mjs obtains it from its declared origin and writes it to the target
project — see Env Vars → The setter (#443). The ship
chain’s stepEnv calls it automatically when the gate is red.
Rule: a red gate blocks or is deleted — no third state. A gate that cannot run prints
WARN: <gate> did not run — <why> (the hook output, the CI step summary, and release-cut.mjs’s
manifest, which now carries a GATES THAT DID NOT RUN line instead of a silent skip).
release-cut.mjs passes --prod-verify-strict unconditionally since 2026.36. Budget: < 60 s locally,
no network (measured 3–5 s warm) — except selftests, which costs minutes by design and carries its
own budget, so the report never has to lie about the 60 s one. Proof harness:
node scripts/agents/landing-gate-selftest.mjs (41 checks: the 2026.35 red list green, planted
violations red, a real git push through the hook, the gate-source override, and group S — the
selftests check in both directions).
A gate is never a function of the branch it judges (#430, #447). Process v4 cuts every work branch
from main, so the gate tooling sitting next to it is last release’s — and a gate younger than the last
ship is simply absent (git skips a core.hooksPath that does not resolve, in silence, exit 0). Gates
are therefore resolved from the shipping side:
scripts/agents/gate-run.mjs— the pre-push entry point. It materialisesorigin/dev’sscripts/agents+scripts/githooksinto.git/ttn-gates/snap-<treesha>/(viagit show, no network) and runs THATlanding-gate.mjs, printinglanding-gate: from origin/dev @ <sha12>. The snapshot is built in a private temp dir and renamed into place (#524), and it counts as cached only when every filels-tree origin/devlists is on disk — a stamp alone is not trusted. Before #524 it was written in place; two concurrent pushes left 37 of 83 files under a.completestamp, the runner reportedlanding-gate.mjs is not present(false) and two release-train pushes went through ungated. A gate that cannot run is a WARN on a work branch and a REFUSAL on a push todev,release/*ormain; an “absent” claim is checked withgit cat-file -ebefore it is printed.- The child gates come from the same snapshot (#447):
gate-run.mjspasses--agents-dir <snapshot>/agents(env spellingTRYVIO_GATES_DIR), andlanding-gate.mjsspawnsproof-lint.mjs/real-seam-gate.mjsfrom there, printinglanding-gate: child tools from <dir>. Without it amain-based branch had its proof judged by LAST release’s linter, so a proof written to today’sWaived:standard (#389) read RED — #424 had to carry every waiver in both spellings, and #430 reproduced the same trap on itself. With no override (CI, where the checkout IS the target, or a hand-run in a checkout) the tree’s own copies are used; a tool missing from the override falls back to the tree and says so — never a silent skip. land.mjsruns the merged tree’s copies — the target plus this branch, i.e. the rulesdevlives by the second the landing publishes. It is the authoritative gate.land.mjsitself comes from the same snapshot (#507). Run from a work branch, its first act is to re-executeorigin/dev’sland.mjsout of.git/ttn-gates/snap-<treesha>/agents(argv, env and the exit code pass through; the child carriesTRYVIO_LAND_REEXEC=1, so it cannot loop) and it printsland.mjs: from origin/dev @ <sha12> (land.mjs <blob>; this tree's copy <blob> is OLDER|NEWER …). Without it every console landed with LAST release’s landing tool, so aland.mjsfix reached nobody until it shipped. No snapshot → the tree’s own copy runs behind aWARNnaming both versions.TRYVIO_LAND_REEXEC=localruns the tree’s copy on purpose (trying aland.mjschange). A branch only re-execs once its own copy carries this code, i.e. from the release that ships #507 on. Regression:node scripts/agents/land-reexec-selftest.mjs.- Audit/repair every existing worktree:
node scripts/agents/gates-doctor.mjs [--check|--apply]. Regression:node scripts/agents/gate-source-selftest.mjs(21 checks, realgit pushes in a throwaway repo with a bare origin).
Why selftests exists (#446). land-selftest.mjs — the proof net for the WHOLE Process-v4 release
chain — was 46/53 on main and 66/73 on dev, red since #401/#399 landed, and nobody found out,
because nothing ran the family: no hook, no workflow, no gate. #455 measured four more red the same way,
all of them stale assertions rather than broken tools. On ship day that makes a real regression in
release-close.mjs indistinguishable from months of its own noise. The register is linted against the
directory precisely so the failure cannot recur: add a harness without declaring it and the gate is RED
with the line to paste. The register is read from the JUDGED tree, so a push that adds a harness and
declares it is not judged by origin/dev’s older register.
proof-lint ergonomics (#389). Filename ranges ac1-stage-1..5-after.png and brace lists
ac2-{light,dark}-390.png are expanded and every member must exist. Waivers are first-class:
Waived: Harness — <why> · Waived: Screenshots — <why> · Waived: 390px — <why> ·
Waived: Mock-mode — <why> (legacy Harness: N/A — … / Screenshots: N/A — … still count); a
waived check passes AS WAIVED and is listed distinctly for owner audit. Mock mode: is recognised
in bullet/bold form (- **Mock mode: on**). --range resolves issues from declared refs only
(_issue-refs.mjs, #360). The PROOF.md template in 05_tasks/proof/README.md carries every
lintable stamp as a placeholder.
Transition. During release 2026.36 the old cherry-pick release-cut.mjs path (below) stays
runnable; it retires once both paths produce the same tree.
Proof harnesses. node scripts/agents/land-selftest.mjs (46 checks in a throwaway repo,
including a real publish race) and node scripts/agents/test-branch-guard.mjs (the branch-guard
denial cases).
Release gates (release-cut.mjs, Process v3)
The weekly release assembly (scripts/agents/release-cut.mjs) cherry-picks each approved issue’s
dev commits onto release/<year>.<week> and runs a fixed gate sequence before the owner’s single
Wednesday “шипвай”: conflict-stop (never auto-resolve) → content-equivalence vs origin/dev →
widget-asset rebuild → proof-lint → real-seam + dry-run (money/DB issues, #263) → tsc + vitest.
Skipped under TRYVIO_RELEASE_SELFTEST=1 (used by release-selftest.mjs).
Real-seam + dry-run gate (#263)
For every manifest issue labeled zone:billing or zone:db, the gate runs
scripts/agents/real-seam-gate.mjs scoped to that issue’s changed files. A money/DB diff
(02_app/apps/web/src/lib/billing/**, supabase/migrations/**, or an added .rpc(/SQL line in
server or API code) fails the release unless one of:
- a committed
scripts/real-seam/rs-<issue>-*.mjstest exists (or the diff itself adds one), or - the issue’s
PROOF.mdcarries a literalReal seam: N/A — <why>waiver (printed loudly in the gate output).
With --money (set automatically for zone:billing issues) the issue’s PROOF.md must also carry
a literal Dry-run: block — see Billing pipeline → Dry-run mode.
Real-seam tests are plain node scripts in scripts/real-seam/ (shared plumbing _lib.mjs) that
execute the actual SQL/RPC against the prod-shaped dev Supabase via the Management API — the dev
project ref is hard-coded, so a real-seam test can never point at prod. They exist because vitest
mocks exactly the seam that broke prod twice: #219 refunded $0 of 367 units because the
LIMIT-before-predicate bug was invisible against dev’s tiny ledger, and #246 mocked is_test away
entirely. Reference test: scripts/real-seam/rs-263-no-job-sweep-reach.mjs — the old query shape
finds 0 candidates, the real no_job_sweep_candidates() RPC finds 367 on the same seeded ledger.
The seeded data underneath. node scripts/agents/seed-prod-shape.mjs --confirm=SEED-DEV (hold
the agents’ db-migrate lock while it runs) creates two synthetic stores on dev Supabase:
seed-icedout-shape— traffic/money worst case: 10ktryon_billing_ledgerrows carrying the #219 adversarial distribution (~9.6k delivered rows OLDER than the 367 refundable ones), 17.6k sessions, 100k analytics events.seed-fatcatalog-shape— catalog worst case: 3,064 products / 12k images.
Plus prod-order rows for four tables that were otherwise empty on dev: attributed_orders (570),
attributed_order_lines (820), captured_emails (178), shopper_demand_signals (57). UUIDs are
deterministic (md5-derived), so the seed is idempotent; --wipe deletes exactly the two seed
stores’ rows in FK order; --verify checks 24 per-table targets.
proof-lint: mock-mode stamp now required
proof-lint now errors (was warn-only) when a PROOF.md claims provider behaviour
(kie/fal/generation) or DB behaviour (ledger/rpc/sql/migration/refund) without a literal
Mock mode: <state> line — the #107 “13/13 всичко бачка” proof had run fully mocked, and the
stamp is what makes the mode visible in the release manifest review.
Post-deploy verification (prod-smoke.mjs, #345)
A green deploy proves the code shipped — not that the features work. v2026.34 was green on every
gate (tsc, vitest, proof-lint, content-equivalence, schema-drift, health 200, the real-shopper
try-on smoke) and still shipped the affiliate portal dark (#343) with an unknown-email
dead-end (#344); the owner found both in his browser. scripts/agents/prod-smoke.mjs is the
codified answer — one read-only command, a PASS/FAIL table, non-zero exit on any failure.
node scripts/agents/prod-smoke.mjs --target=prod # release-agnostic subset
node scripts/agents/prod-smoke.mjs --target=prod --issues 343,344 --telegramTwo layers:
- Release-agnostic — always true of a healthy prod, and what CI runs after every ship:
/api/health/deep200 with every componentokand thefeature_configper-featureconfiguredbooleans (#343 — a dark feature is invisible to a status check by design, because health deliberately staysokfor it); public surfaces (/affiliate,/affiliate/apply,/api/health, landing, docs) returning 200; the live widget build read off the storefront’s CDN asset (version +--markerscontent greps — the only honest proof of which build is live); and theapp_logserror window (--since-minutes, budget--max-errors, default 0). - Release-declared — each shipped issue’s
Prod-verify:lines from itsPROOF.md(GET/SQL/CDN marker/manual:; convention in05_tasks/proof/README.md).release-cut.mjscollects them into the manifest’s VERIFY ON PROD checklist and names the command in the ship steps;proof-lintwarns when an issue declares none (errors under--prod-verify-strict). Checks needing an embedded Shopify admin session (STANDARDS §9.8) are declaredmanual:with an assignee and reported as MANUAL — never auto-passed. These declarations are now actually run, with evidence, byprod-verify.mjsbelow; CI runs their automated half on every deploy.
Invariants: read-only (only GETs and single SELECT statements — enforced in code, not by
convention) and fail-closed — an unreachable API, a missing credential or an unparseable
declaration is a FAIL naming the piece, never a silent pass. Budget 90s; measured ~1s.
production.yml runs the release-agnostic subset after the real-shopper try-on smoke with
--since-minutes=5 (a wider window looks back past the deploy and would blame the new release for
older errors), skipping loudly when its credentials are absent. Credentials: SUPABASE_ACCESS_TOKEN
plus SMOKE_STORE_PASSWORD/TESTSTORE_PASSWORD (or ~/.tryvio/smoke-store-password).
Post-ship evidence pack (prod-verify.mjs, #401) — a ship is done only when prod is PROVEN
Owner rule (2026-08-20, on #387): after a prod ship, smoke tests plus everything an agent can test
on prod must actually be tested — screenshotted, recorded, or called — and the evidence sent; the
ship is not done until prod is proven. v2026.34 shipped the affiliate portal dark and the owner
found it; the 2026.35 manifest assigned 25 manual: owner checks — work handed to the owner
instead of results.
node scripts/agents/prod-verify.mjs --release 2026.35 --telegram
node scripts/agents/prod-verify.mjs --release 2026.35 --finalize --telegram--release <year>.<week>[.<patch>] (a .patch tag scopes the pack to that
hotfix issue); --issues a,b overrides the default scope.
The default scope is the release itself, not a tag range (#455). A train week resolves to its
land(N) promotions — before the merge from the open branch
(origin/main..origin/release/<w>), after it from the release merge commit
(<merge>^1..<merge>, i.e. the release side and nothing else); _release-scope.mjs
answers both. Only when a release has neither branch nor merge commit — a hotfix tag, --ci with no
week — does it fall back to <previous tag>..origin/main and scan commit subjects
(conventional scopes fix(358):, #id, Refs: trailers), and the scope note says so. The fallback
used to be the default: on 2026.36 it took v2026.35.1..origin/main — 111 commits — and verified
eleven issues from earlier releases, producing 21 FAIL rows that had nothing to do with the release.
deploy-local.mjs --week <w> now passes the week and issue list through, so the fallback
console gets the same scope CI does.
It resolves the Vercel deployment actually serving prod (VERCEL_TOKEN env or
~/.tryvio/vercel-token), then runs prod-smoke.mjs --target=prod --issues … --json --evidence-dir <pack> — health, public surfaces, the live widget build, the error window,
and every shipped issue’s Prod-verify: declaration — retrying a red run up to 3× and reporting
any flap rather than hiding it. --evidence-dir (new in #401) makes every check write its raw
evidence (the GET + status + body excerpt; the SQL + rows; the asset + marker context) into the
pack. It then builds 05_tasks/proof/release-<w>/prod/PROD-VERIFY.md (columns: issue ·
check · evidence · verdict · timestamp · deployment) plus prod-smoke.json, pack.json, per-issue
folders <issue>/ (auto evidence get-N.txt, sql-N.txt, cdn-N.txt), and _release/
(release-agnostic evidence).
Agent evidence, not owner homework. Every issue whose proof was UI-flavoured gets a “UI
evidence — prod screenshots (both themes where the surface has them)” row; widget-flavoured gets a
“widget evidence — recording on the prod smoke store tryvio-pgi2n05r” row; every manual:
declaration gets a row too. These start ⏳ TODO until a console writes <pack>/<issue>/EVIDENCE.md:
Evidence: <file-or-url> — <what it shows> -> PASS|FAIL
Cannot: <check> — <why an agent cannot do it on prod>; proxy: <read-only thing checked instead> -> PASS|FAIL|NONE
Waive: <check substring> — <why the automated red is false>; evidence: <file> -> PASSWaive: turns an automated FAIL the agent can SHOW is false (e.g. a 1 day SQL window that straddles
the deploy — the 2026.35.1 pack’s case) into ⚠️ WAIVED: never PASS, listed in its own section and in
Telegram; the evidence file must exist.
manual: owner is never the default — an owner-assigned declaration stays TODO until an agent
covers it with an Evidence: line or an explicit Cannot: line naming the read-only proxy it
checked instead (listed under “Cannot be checked on prod by an agent”). --widget-video
optionally runs the real shopper try-on harness (playwright-debug/prod-smoke-tryon.js) on the
smoke store and copies the recording into _release/ (one real generation; opt-in).
Status line (parsed by the gates): Status: INCOMPLETE — n TODO / Status: COMPLETE GREEN /
Status: COMPLETE RED. --finalize refuses (exit 1) while INCOMPLETE; once complete it sends one
Telegram message (verdict, deployment, counts, every non-green row, rollback line) with up to 12
shots + 3 videos, stamps Finalized: <ts> · Telegram: sent (n photo(s), m video(s)) from the
notifier’s actual return value (a failed send is stamped NOT sent, exit 1, and release-close
refuses — the owner must HAVE the pack), and publishes the pack to dev via a temp worktree
(proof(401): release-<w> post-ship evidence pack — …; the release folder gets a minimal
PROOF.md when it has none — landing-gate lints every touched proof folder); --no-publish/--dry-run
skip that. RED prints the rollback command
(prod-rollback.mjs --release=<w> --confirm=ROLLBACK-PROD --telegram), and every failed
issue is re-opened + labelled keep-open (board-sync won’t auto-close it again) with a comment
linking the evidence.
Gates that read the pack. release-close.mjs --week <w> aborts unless the pack (origin/dev
→ origin/main → local tracked) shows Status: COMPLETE GREEN — no pack, no tag, no back-merge.
proof-lint.mjs --prod-pack <rel> checks the pack exists, is COMPLETE, carries the deployment
id, and has ≥1 tracked evidence file per issue it names; release-cut.mjs --freeze runs it for the
previous ship (newest tag) — ships tagged before 2026-08-22 are grandfathered with a WARN. The
release manifest’s VERIFY ON PROD section is now the pack’s own checklist plus both prod-verify
commands.
CI. production.yml’s post-deploy step now runs prod-verify.mjs --ci --since-minutes=5
(checkout with fetch-depth: 0) — the automated half only (release-agnostic checks + every shipped
issue’s Prod-verify: declarations); red fails the deploy job; it writes nothing to the repo. The pack
itself is finalized by a console. Since #716 the CI run scopes to the release it deployed — the week is
read off the release merge at HEAD, so the issues are that release’s land(N) promotions, exactly as
the console run; newest tag..HEAD stays the fallback for a head that is not a release merge (a
hotfix). And because CI runs right after the API deploy — the widget deploys after the API — a missing
CDN marker is PENDING there (CANNOT, not a FAIL); the console run after the widget deploy enforces it.
The ship’s deploy step waits for a production run that has STARTED until it completes (cap 60 min;
the job’s own timeout is 40): a still-running run is a wait, never “a real failure” (the 2026.39 stop).
Other flags: --smoke-json <file>, --deployment <id>, --dry-run, --attempts N.
Proof harness: node scripts/agents/prod-verify-selftest.mjs (offline, throwaway repo:
INCOMPLETE→refuse, phantom evidence→FAIL, COMPLETE RED/GREEN, proof-lint --prod-pack,
release-close gate, --ci). Convention: 05_tasks/proof/README.md § “Post-ship evidence pack
(#401)”; law: 05_tasks/STANDARDS.md §7 and §10.
Touching the production database (#451) — one gated door, fail-closed
No console, script or local process reaches the prod database without the owner approving THAT
operation first. Owner rule 2026-08-24; the DB sibling of golden rule #1, because a bad write is not
undone by re-pointing a deployment. It exists because on 2026-08-24 a console copied
02_app/apps/web/.env.local (which is prod) into a work tree and ran a local Next dev server
against it — harmless by luck, and caught by reading the env afterwards, not by anything refusing.
The choke point
scripts/agents/_prod-db.mjs is the only module allowed to name the prod Supabase ref, and it owns
the transport, the read/write classifier, the grants and the audit log. prod-state.mjs re-exports its
supabaseQuery as supabase, so prod-snapshot, prod-verify, the real-seam helpers and the ship
chain’s migration apply are all gated without a change at their call sites. schema-drift,
prod-smoke, placement-baseline and manual-block-usage import it directly.
node scripts/agents/prod-ref-lint.mjs— RED if any file underscripts/,.claude/,.github/or02_app/can reach the database outside that one module. Archived scripts under05_tasks/proof/**are reported, not enforced (they are the committed evidence of finished issues; two of them can still write to prod — do not run them).- Mention ≠ reach, and the distinction is structural, not a path allow-list.
/storage/v1/object/public/<bucket>/<object>is the one Supabase route family served anonymously — an outreach email’s<img>src on that route is a CDN-class link and is listed underPUBLIC ASSETS, not as a violation. Everything else on the project origin (/rest/v1,/auth/v1,/functions/v1,/realtime/v1, and/storage/v1/object/…withoutpublic/, including thesign/route that mints private-bucket URLs with the service-role key) needs a credential, as doesapi.supabase.com/v1/projects/<ref>— those are reach. Default is reach: a bare ref constant, an unparseable token, a public route with no object, and a “public” URL carrying?token=/?apikey=all count. One definition (classifyProdRefUse), shared by the lint and the branch-guard rule below. - The gate lives in the code path, not in a hook:
branch-guardis a PreToolUse hook on a typed command string and is blind to a spawned process. A child re-enters the module and re-checks, because a grant is a file; a claimed write batch lives in process memory, so a child cannot ride on its parent’s approval. .claude/hooks/branch-guard.mjsrule 8 covers what the code gate cannot see — a hand-typedcurl/node -e/psqlat the prod project is denied and pointed here.
Asking for an operation
# a read — cheap, reusable by that one tool, 60 minutes
node scripts/agents/prod-db-approve.mjs read --tool schema-drift.mjs \
--reason "post-deploy schema gate" --quote "<the owner's verbatim words>"
# a write — names the exact statements, bound to their SHA-256, single use, 15 minutes
node scripts/agents/prod-db-approve.mjs write --tool fix.mjs --sql-file fix.sql \
--reason "one stuck row on #NNN" --quote "<the owner's verbatim words>"
node scripts/agents/prod-db-approve.mjs status # what is live right nowGrants live beside the PR approval nonces in ~/.tryvio/approvals/ and are burned by the same rename,
so “used” means the same thing everywhere. A read grant never authorises a write, and editing the
SQL after the yes changes the fingerprint and voids the grant. Yesterday’s approval is not today’s.
Refusal is exit 3, with the exact command to ask printed. An internal error in the gate refuses too — unlike the fail-open Claude hooks, this one sits in front of data.
Read vs write
A statement is a read only if it is a single select/with/explain/show/table with no
mutating keyword and no function call outside the read-safe list. Everything else is a write —
including select analytics_rollup_rebuild(...), which deletes and inserts. A false “write” costs one
approval prompt; a false “read” costs an unapproved mutation.
What still writes to prod, and under what approval
Exactly one thing: the release’s migration apply in _ship-steps.stepMigrations. The owner already
approved that release by name with its migrations listed in the ONE SCREEN, so the chain derives its
grants from the шипвай nonce (ensureShipDbAccess for reads, a statement-bound write grant for the
DDL) instead of stopping to ask a second time. The N migration files plus
notify pgrst, 'reload schema'; are ONE claimed batch — a statement that was not in it still refuses.
The PR nonce is read, never burned there: stepMerge still owns the burn.
deploy-local.mjs --target=prod --confirm=SHIP-PROD derives a read-only grant from the typed
confirmation for its post-deploy battery.
After the merge, the nonce is burned — but health, schema-drift, prod-smoke and the evidence pack
still read prod, and they are verifying the release that nonce just authorised. So stepMerge opens a
read-only 6-hour tail as it burns. That is what keeps ship.mjs --from prod-verify and a later
prod-verify --finalize from becoming a new ship-day step. A write after the merge still needs its own
statement-bound approval, and a pack finalised days later is not ship day any more: it asks for a read
grant the ordinary way (the refusal says so, and explicitly tells you not to re-approve a PR that
is already merged).
CI, local apps, and the escape hatch
- CI (
production.yml’s post-deployschema-drift/prod-verify --ci) may read, with the repo and run id in every audit line, and may never write — a runner cannot approve itself. - A local app cannot be pointed at prod:
apps/web/next.config.mjsrefuses a localnext dev/next buildwhose resolvedSUPABASE_URLis not the dev project, and names the file that set it. It is inert on Vercel, in CI and under vitest — the deployed app obviously talks to its own database on every request, and this issue is about us, not about the product. - Break-glass is a flag, never an env var:
--break-glass-prod-db="<why>"lets the operation through, prints a banner, writes abreak-glassaudit row and sends Telegram. SettingTRYVIO_PRODDB_SKIP/_ALLOW/TRYVIO_SKIP_PRODDB_GUARDdoes not disable the gate — it refuses louder.
Audit
Every grant, read, write, batch claim and break-glass appends one JSON line to
~/.tryvio/prod-db-audit.log (outside the repo — statements can carry shop domains and ids; tokens
never appear). An unwritable log refuses the operation. Telegram is sent for grants, writes and
break-glass — the events that change what prod can become — not for each read inside a granted
session, which is what the log is for.
Honest limitation: this binds the repo’s tooling. It cannot stop a raw curl carrying the
operator’s own Supabase token from a shell nothing observes. Closing that needs a read-only prod
credential, not a lint.
Harness: node scripts/agents/prod-db-selftest.mjs (hermetic — temp approvals dir, no network, no
prod, no Telegram).
Migrations
SQL migrations live in apps/web/supabase/migrations/ and are applied manually, by the ship chain,
through the gate above (stepMigrations) — never by hand from a console without an approval.
Adopting the Supabase CLI (supabase db push) is the first task of the infra roadmap; whatever
replaces the manual apply has to go through _prod-db.mjs too.