Architecture

Architecture

High-level

Shopper (storefront)                    Merchant (Shopify admin)
      │                                          │
  Theme app extension                    Embedded app (App Bridge)
  "Try it on" widget                     /merchant, /settings, /billing …
      │                                          │
      └──────────────► Next.js app (Vercel) ◄────┘
                        app.tryvio.ai
                              │
        ┌─────────────────────┼─────────────────────┐
        ▼                     ▼                     ▼
   Supabase              Shopify Admin API      AI providers
 (Postgres + Storage)   (GraphQL, webhooks,    (KIE.AI / fal.ai)
                          Billing, Web Pixel)

Surfaces

  • Storefront widget — a Shopify theme app extension injected into the merchant’s theme. Calls the public storefront API routes (/api/storefront/*) and the app proxy (/api/proxy).
  • Embedded merchant app — Next.js pages rendered inside Shopify admin via App Bridge (/merchant, /settings, /billing, /plan, /products, /brand, /discounts, …). Detected via the x-tryvio-shell: embedded header / host param. /plan is the mandatory-plan gate page — see Billing Pipeline → Mandatory plan gate.
  • Public marketing site — the landing page at / (when no Shopify host/shop), plus /contact, /privacy, /terms, and the standalone live demo.
  • Admin/operator dashboard — /admin, PIN-gated, for internal operations.

Request routing

src/app/page.tsx is the fork: Shopify traffic (has host or shopify=installed&shop) redirects to /merchant; everyone else gets the public <LandingPage />.

Key external services

ServicePurpose
SupabasePostgres (all app data) + Storage (input/output/merchant image buckets)
Shopify Admin APIOAuth, product sync, orders, billing, discounts, Web Pixel
KIE.AI / fal.aiAI try-on image generation
VercelHosting, cron jobs, env management

Cron jobs

Defined in apps/web/vercel.json:

  • /api/cron/billing-overage — daily 02:00, charges pending overage.
  • /api/cron/trial-expiry — daily 02:45, sweeps lapsed/reopened trials into trial_expired/ trial — see Billing Pipeline → Trial lifecycle.
  • /api/cron/trial-comms — hourly (25 * * * *), the 3-day / 1-day trial warnings and the goodbye (in-app + email, once per trial window) — see Billing Pipeline → Trial warnings.
  • /api/cron/review-ask — daily 09:15 UTC, the App Store review ask’s bell (trial over) and email (subscription day 43) — see Admin & landing → The App Store review ask.
  • /api/cron/storage-cleanup — daily 03:00, expires old try-on images.
  • /api/cron/widget-health — every 6h (20 */6 * * *), the widget-health watchdog — see below.
  • /api/cron/activation-chase — hourly (50 * * * *), the activation chase — see below.
  • /api/cron/quality-watchdog — daily 04:10, the try-on satisfaction watchdog — see below.
  • /api/cron/support-sweep — hourly (0 * * * *), remind then close support conversations — see below.
  • /api/cron/copurchase-pairs — daily 03:30, computes co-purchase (bought-together) pairs from each store’s Shopify orders — see Try-on Pipeline → Co-purchase tier.

Support sweep — remind, then close (#324 / #352)

GET /api/cron/support-sweep makes two passes over the OPEN threads, and they never touch the same row. Which pass a thread belongs to is decided by one column: support_threads.last_sender.

  • Remind (#352) — the merchant spoke last and has been waiting longer than FIRST_REMINDER_HOURS (4). The nag goes into the merchant’s own Telegram topic (replying is one tap) and as one line into the ops chat with a t.me/c/… deep link — deliberately a second chat, because a reminder posted only into the topic lands where the owner already scrolled past once. It repeats every REPEAT_REMINDER_HOURS (24) until we answer or the topic is /closed.
  • Close (#324 door B) — nobody has written for INACTIVITY_DAYS (3) and the last word was ours. Merchants rarely say “this is resolved”; they just stop writing.

Why the last_sender split matters: “nobody has written for three days” means two opposite things — the merchant is satisfied, or we never answered — and closing both alike is how the first support ticket we ever received (2026-08-12, a paying merchant reporting #350) would have been filed away with a green tick. It sat unread for 24 hours.

  • State without a migration — the reminder’s bookkeeping lives in support_threads.metadata.unansweredReminder = { forMessageAt, count, lastAt }. It is bound to the message being waited on, so an hourly cron cannot become an hourly alarm, and a merchant who writes again starts a fresh wait instead of inheriting a spent schedule.
  • A failed send is never recorded as a reminder — the next run tries again. Marking a refused Telegram call as “reminded” would rebuild the exact silence the pass exists to break.
  • Decision logic is a pure module — lib/support/thread.ts (shouldRemind, describeWait, telegramTopicLink), unit-tested; the orchestration in lib/server/support-chat.ts (remindUnansweredThreads) only wires it to I/O. The two PostgREST filters that carry the split are covered by scripts/real-seam/rs-352-unanswered-nag.mjs against real dev rows, because a mocked test cannot prove a query.

We speak first — the install / activation greeting (#609)

Two moments send an automatic sender: "owner" message into the merchant’s support thread, before the merchant has ever written anything: app install (/api/internal/setup calls greetMerchant({ store, kind: "install" }) — started before the product sync and awaited after it, so the hello never waits on a big catalogue) and a subscription’s first activation (both /api/shopify/billing/callback, on !alreadyActivated, and the app_subscriptions/update webhook, on first activation, call greetMerchant({ store, kind: "subscription", subscriptionId })). The reconciler’s repair-path activation deliberately does not greet. Both activation paths load the store through the one shared column list STORE_RECORD_COLUMNS (#748): the webhook’s getStoreBySubscriptionId used to select its own shorter list without dashboard_locale, so whenever the webhook won the dedupe race the plan greeting fell back to English (a Bulgarian store, prod 2026-09-30).

greetMerchant (src/lib/server/support-chat.ts) runs claim → store → mirror:

  1. Claim — createNotification with dedupeKey support-greeting:install or support-greeting:subscription:<subscriptionId>. The notifications (store_id, dedupe_key) unique index is the exactly-once gate, per store: a reinstall, a retried setup, or the callback+webhook pair of one activation lose the race and stop (createNotification returns false on a duplicate).
  2. Store — insertSupportMessage with metadata: { greeting: <kind>, locale }.
  3. Mirror — Telegram, best-effort and never throwing: ensureTopic creates the merchant’s forum topic at the greeting (so it exists before the merchant ever writes), then a note 🤖 Auto-greeting (install · ro) — shown to the merchant in the app: &lt;text&gt;. Install/billing never fail over this. Logs: support.greeting.sent, support.greeting.duplicate, support.greeting.failed, support.greeting.locale_fallback (raw + resolved, when dashboard_locale is NULL, unshipped, or region-tagged like pt-BR → pt).

The copy comes from the app i18n catalogue — support.greeting_install and support.greeting_subscription ({plan} = plan name) — rendered for stores.dashboard_locale via the new server helper translatorFor() (src/lib/i18n/translate.ts, no request headers, the same catalogue the dashboard renders). The bell entry title is notifications.support_greeting_title; the old hardcoded-English welcome notification now uses notifications.welcome_title / notifications.welcome_body. All five keys exist in all 24 locales.

Operator messages from the console (#745). The owner approves a text; a console sends it to ONE merchant through all the support channels at once:

node scripts/agents/merchant-message.mjs --store <id|x.myshopify.com> --message-file <txt>   --subject "<email subject>" [--target prod --confirm --quote "<owner's approval>"] [--dry-run]

The tool only carries the text: it calls POST /api/ops/merchant-message (Bearer OPS_API_SECRET, timingSafeEqual, fails closed when unset), and the server runs sendMerchantMessage (src/lib/server/merchant-message.ts) → postOwnerMessage — claim (bell, dedupe_key operator-message:<sha256(text)>, title in the merchant’s dashboard language, link /analytics?support=open) → chat message (metadata.operator) → support email (merchant’s language, “reply in the chat” button) → mirror into the owner’s Telegram topic. It has to be the server: the Resend key and the support sender are sensitive Vercel env that no console can read — a console-side send was tried first and got chat + bell out while the email silently did not go (Fury Clothing, 2026-10-02). OPS_API_SECRET is created by env-set.mjs as a readable (encrypted, not sensitive) key so the tool can fetch it. A retried send of the same text is a duplicate; dryRun: true resolves the store and reports the language and whether an email can go out.

Auto-open is derived, not stored. GET /api/shopify/support/status (same auth, resolveAuthorizedShopForShopifyRequest, store_id-scoped) returns { unread, autoOpen } from getSupportThreadStatus + the pure supportStatusFor (src/lib/support/thread.ts): unread = the last message is ours and newer than merchant_last_read_at; autoOpen = the newest greeting (a message with metadata.greeting) lives in the OPEN thread and is newer than merchant_last_read_at. Opening the panel calls GET /support/messages, which marks the thread read — that consumes the auto-open on every device, with nothing in localStorage. A plain owner reply lights the dot but never auto-opens. The closed bubble polls /support/status every 20 s; the greeting’s bell entry links to /analytics?support=open, which the bubble honours on mount.

needsContact / canStartNew in GET /support/messages now come from conversationFlags (has the merchant written), not messages.length — otherwise a thread holding only our greeting would skip the contact step forever and offer “close” on a conversation the merchant never wrote in. The topic name (composeTopicName) and the pinned header (composeContextMessage) now carry the merchant’s dashboard language (… · starter · ro, plus a 🌐 ro row).

Fixed on the way: getStoreByDomain never selected dashboard_locale, so the “localised” welcome email (#476) went out in English to every merchant; the column is now in that SELECT.

The reply email leads back into the chat (#457). When the owner answers from the topic and the merchant is not reading the thread (shouldNotifyMerchant), notifyMerchantOfReply emails the reply (sendSupportReply). Owner decision 2026-09-21: support runs through the chat, a mailbox is not ingested into threads — so that letter now carries a Reply in the chat button to the embedded app (/analytics?support=open) and one line saying the chat is where we answer fastest, in the merchant’s dashboard_locale (getStoreById now selects it; the letter used to go out in English). Without it the merchant’s natural move is Reply, which lands in a mailbox and not in the thread. A reply sent by email anyway is still caught by the morning inbox sweep (#518). An admin reply sent to a bare address (/api/admin/support/reply) has no store behind it and stays a plain letter.

Known limitation — the inactivity sweep (closeStaleThreads, 3 days, above) still closes a thread whose only message is an unanswered greeting; a merchant who first opens the dashboard more than 3 days after install finds the greeting under “Previous conversations” and in the bell, not as a popup.

Try-on satisfaction watchdog (#305)

GET /api/cron/quality-watchdog is the answer to “36 % of rated try-ons are 👎 and nothing notices”. Once a day it scores every real store’s trailing-7-day dissatisfaction and sends the owner one Telegram per slump. Same skeleton as the widget-health watchdog: isAuthorizedCronRequest, acquireCronLock("satisfaction_watch_lock", 10), one set-based RPC, all judgement in a pure lib.

  • The detector — store_satisfaction_stats(p_window_from, p_baseline_from, p_to) (20260812c_store_satisfaction_stats.sql, service_role-only) returns every real store with its 7-day counts, the distinct shoppers behind them (via tryon_sessions.metadata->>'shopperPseudoId'), its own 30-day baseline, and the top affected product_type. One pass over the baseline slice; the window is a FILTER inside it, not a second scan.
  • The thresholds are measurements, not taste (prod, 2026-08-12 — full derivation in 05_tasks/proof/305/PROOF.md). They live as exported constants in lib/quality/satisfaction-watchdog.ts, and the admin chart imports them so the screen can never disagree with what fires:
    • absolute bar 40 % — the fleet baseline IS ~36 %, so a 35 % bar fires on an ordinary week;
    • relative rule +10 pp over the store’s own 30-day rate — catches a 15 % store that doubles to 27 %, which no absolute bar would ever see;
    • floors ≥20 rated AND ≥8 distinct shoppers — the real prod tail reads 66.7 % on 3 ratings, and the shopper floor is what stops one angry pair from tripping an alert;
    • 5 pp hysteresis so a store parked at 40.2 % cannot flap.
  • Classes: no_data (nobody rated — not the same as 0 %), insufficient_sample, concentrated (a burst, not a signal), healthy, elevated (absolute), regressed (relative).
  • Episodes — state in app_config.satisfaction_watch_state. One alert per episode; an episode closes only on a scored window back under the bar. Silence is never treated as recovery.
  • Operator surface — /admin?tab=quality shows dissatisfaction over time, sliceable by model and by category, from admin_satisfaction_timeseries (one read backs the chart and both slicers). Model attribution is only truthful after #303’s backfill; unrecoverable rows render (unattributed) rather than averaging a guess in with the facts.
  • No merchant comms here on purpose — telling a merchant “a third of your shoppers dislike the results” is a conversation, not a notification. #290 owns the merchant-facing readout.

Widget-health watchdog (#310)

GET /api/cron/widget-health finds stores whose widget has gone invisible: shoppers are landing on product pages, but the widget never renders — a broken or never-enabled theme app embed. Auth is the shared isAuthorizedCronRequest (timing-safe CRON_SECRET bearer, same as the other cron routes); an overlap guard, acquireCronLock("widget_health_lock", 10), keeps two overlapping runs from racing the episode state below.

  • One set-based detector, no per-store fan-out. widget_health_stats(p_from, p_to) (migration 20260811_widget_health_stats.sql, service_role-only) returns every real store (is_test = false, not uninstalled) with its windowed product_viewed/impression counts and entitlement fields in a single query.
  • Classification (lib/widget/health-watchdog.ts, pure + unit-tested) looks at the trailing 48h for every store with ≥1 synced product and ≥1 product_viewed event:
    • healthy — impressions in the window;
    • no_catalogue — nothing synced yet (#312’s wall, not this watchdog’s fault);
    • too_new — installed inside the 48h window, no fair verdict yet (#238 owns day one);
    • quiet — no shopper views at all — never flagged (absence of traffic ≠ absence of the widget);
    • paused_entitlement — resolveStoreGate (lib/billing/store-gate.ts) — the same gate the storefront widget-access path uses — says the store is gated: trial expired/exhausted, cancelled, or frozen. Dark by design; owned by #311, never alerted here;
    • paused_merchant_off — widget_enabled_globally = false or product pages turned off;
    • invisible — synced catalogue + real shopper traffic + zero impressions + we SHOULD be serving. The only class this cron acts on.
  • Episode state, not a stateless sweep. app_config key widget_health_state holds a JSON map storeId → { since, alertedAt, escalations? }. An episode closes when impressions resume, when the store leaves the sweep (uninstalled / marked test), or — since #525 — when the storefront gate (resolveStoreGate) pauses it. The gate is asked directly, not via the class, so an expired trial that has also gone quiet closes too. An expired trial is dark by design, so chasing it about its theme embed would be false (3 of the 5 episodes open on 2026-08-31 were expired trials that could never close). A store that merely goes quiet keeps its episode — no traffic is no verdict. A relapse after a close starts a fresh episode.
  • The ladder (#525). Until #525 an episode alerted once, on day 0, and never again — on 2026-08-31 five were open, the oldest 19 days, one of them a paying Growth account. Each open episode now climbs, one step per run at most, and only while the store is still invisible:
    • day 0 — notice (the original alert);
    • day 3 — reminder (“still nothing”);
    • day 7 — final, in the voice of what the outage costs: a paying store (“you’re paying for a widget nobody can see”), a trial (“your free trial ends on …”), or neither (the reminder’s words). The final says it is the last email; after it the merchant is not written to again.
    • A trial with ≤ 3 days left jumps straight to the final, whatever the episode’s age; a store that turns invisible that close to its trial end opens at the final (one message, not two).
    • Paying vs trial is billingSegment (lib/admin/customer-segment.ts, #724) — the same rule the admin’s segments use. Steps sent are recorded in escalations (step, at, reason age | trial_deadline) before the send, like alertedAt.
    • Backfill: an episode opened before #525 has no escalations; its since is day 0, so one older than 7 days gets the final on the first run after the ship (the reminder is skipped).
  • Alerts (newlyFlagged + escalated of a run):
    • Merchant — in-app notification type widget_invisible (links to /widget) via notifyWidgetInvisible (lib/server/merchant-notifications.ts), plus email (templates.widgetInvisible, step-aware) whose CTA is the absolute Theme-Editor app-embed deep link from buildThemeEditorLinks (lib/server/shopify.ts). Dedupe key per step against the notifications_dedupe_idx unique index: widget-invisible-<episode start date> for the notice (unchanged, so a pre-#525 episode is never re-noticed), …-reminder, …-final. The email fires only when the notification was freshly created. The email is in the merchant’s dashboard_locale — the RPC does not return it, so the route reads it for the alerted stores (getStoreDashboardLocales); unknown degrades to English. The in-app notice is English, like every dashboard notice.
    • Operator — ONE Telegram per run (Bulgarian, shop domains only — no email addresses): the newly invisible stores, the ones chased again, and on the 06:20 UTC run only the backlog — every episode older than 72h with its age, paying/trial state, plan, days to trial end and last step, paying stores first (widgetInvisibleBacklog). No new, no chase, no backlog → no message.
    • Admin — the merchants list shows a red Widget not showing · Nd pill next to the customer segment for every store with an open episode (widgetInvisibleSince on the getAdminMetrics row, read from the same app_config key; a failed read hides the pill, never the list).
  • Dev is inert by env, not by a flag: no RESEND_API_KEY and no TELEGRAM_BOT_TOKEN/TELEGRAM_OPS_CHAT_ID on the dev Vercel project, so both senders no-op gracefully.
  • Calibrated against prod (2026-08-11, read-only): 5 stores were dark-with-traffic; all 5 were correctly classified paused_entitlement (expired trials) — zero false alerts.

Activation chase (#312)

GET /api/cron/activation-chase is the bounded safety net behind install-time auto-sync: the 2026-08-10 funnel audit found 9 of 24 merchants installed, never synced a single product, and were never contacted again. Auth is the shared isAuthorizedCronRequest (timing-safe CRON_SECRET bearer); an overlap guard, acquireCronLock("activation_chase_lock", 10), keeps two overlapping runs from double-retrying or double-emailing. maxDuration = 300, fetchCache = "force-no-store" (the same #97 stale-fetch-cache trap as widget-health — force-dynamic alone doesn’t bypass it).

  • Candidates — listActivationCandidates (lib/server/supabase-admin.ts) finds ACTIVE stores whose latest install/reinstall event (install_events, see Install Attribution) falls between 24h and 14 days old and that still have zero catalog_products. A reinstall’s latest event resets the clock — a fresh cycle.
  • Retry before emailing. For each due store the route first retries the server-side product sync itself (runShopifyProductSync), capped at 2 retries per run — extra due stores are deferred to the next hourly run rather than dropped (their window is 24h–14d wide, so nobody is silently skipped). A store with a sync genuinely running is skipped entirely this run (never starts a second sync over it — the same guarantee as AC5 below).
  • Reason is diagnosed from the retry outcome (pure logic in lib/activation/activation-chase.ts, reasonAfterRetry/planChase, unit-tested):
    • synced_now — the retry brought products in; the next concrete step is the app embed;
    • empty_catalog — sync is fine, the Shopify store genuinely has no products;
    • reconnect — the OAuth token is dead; only reopening the app from Shopify admin fixes it;
    • retry_failed — sync keeps failing (throttle/timeout) — merchant is pointed at Products → Retry.
  • Two attempts, ever. dueChaseAttempt fires at 24h and again at 72h, then never again — not a drip. Merchant notification: in-app type activation_chase (notifyActivationChase, lib/server/merchant-notifications.ts) plus email (templates.activationChase), deduped on activation-<attempt>-<install date> (notifications_dedupe_idx) so a reinstall months later — a fresh cycle — earns a fresh chase while each attempt within one cycle fires exactly once however often the cron runs. Every email send leaves an email_events trace (requested/send_failed, type activation_chase).
  • Operator alert — one Telegram message per run via sendOpsDigestToTelegram, listing every freshly-chased store, its attempt number, and the diagnosed reason; silent when nothing fired.
  • Decision logic is a pure module. lib/activation/activation-chase.ts (unit-tested) owns every rule above; the route itself is thin: auth → lock → listActivationCandidates → plan → notify.

Catalog sync (Shopify → Tryvio)

runShopifyProductSync(shop) (lib/server/shopify.ts) is the ONE full-catalog importer. It is triggered at install (/api/internal/setup), by the merchant’s “Refresh products” and by the /products auto-retry — all through POST /api/shopify/sync. Incremental changes arrive via the products/* webhooks (see Webhooks).

Resilience rules (issue #205 — two real installs onboarded into an EMPTY catalog because Shopify answered Throttled and nothing retried):

  • Throttle retry. Every shopifyGraphql call is wrapped in withShopifyThrottleRetry (lib/server/shopify-throttle.ts): capped exponential backoff (4 attempts, 500 ms → 5 s cap), raised to the bucket-refill time when extensions.cost.throttleStatus.restoreRate is present. A throttled query never executed, so replaying it is safe — mutations included. Only throttles are replayed; any other error surfaces immediately. Retries log shopify.graphql.throttle_retry.
  • Single-flight. Concurrent syncs for the same shop share one run (lib/single-flight.ts), so a manual Refresh landing during the auto-retry cannot start a second full fetch.
  • The sync_jobs row is the status flag (no extra column). job_type = 'shopify_products_full_sync'; getLatestProductSyncJob(storeId) reads it — filtered by job type, because Klaviyo writes sync_jobs too. resolveProductSyncState (lib/products-sync-state.ts) maps it to never | running | failed | completed, and downgrades a running job older than 30 min (dead lambda) to failed so the UI can never spin forever.
  • GET /api/shopify/products returns an additive sync field with that state. /products turns it into an honest empty state (syncing · sync failed + Retry · genuinely 0 products · never synced) and auto-re-syncs ONCE per mount when the recorded sync failed — the merchant never has to click.
  • Cache invalidation. The 60 s product-list unstable_cache is tagged merchant-products:<storeId> (lib/cache-tags.ts); the sync and install-setup routes call revalidateTag on success, so a completed sync is visible on the very next GET instead of up to 60 s later.
  • The ProductSync query also selects combinedListingRole, and follows up with one combinedListing.combinedListingChildren query per PARENT node — this feeds product_groups (sibling products sold as separate colour listings). See Product Groups.

First-sync hardening (#312)

Four more guarantees, shipped alongside the activation chase above, close the gaps that let a store sit at zero products silently:

  • AC1 — the OAuth callback’s background setup can’t be dropped mid-flight. /api/auth/callback fires the install-time sync as a self-fetch to /api/internal/setup (a new, independent serverless invocation), which now rides waitUntil from @vercel/functions instead of being a bare fire-and-forget — a lambda frozen right after the redirect returns could previously drop the fetch before it was ever dispatched, silently skipping the first sync.
  • AC4 — a failed first sync alerts the operator immediately. If runShopifyProductSync throws inside /api/internal/setup, it now also sends one sendOpsDigestToTelegram message naming the shop and the error, instead of waiting for the 24h activation chase to notice.
  • AC5 — POST /api/shopify/sync cannot start a second full sync over a running one. It first reads the store’s latest sync_jobs row via resolveProductSyncState; if a fresh running job exists, it returns { alreadyRunning: true, sync } instead of kicking off a duplicate fetch. A running row older than 30 min is still resolved to failed and re-synced as before — this only guards the cross-invocation race, not a dead lambda.
  • AC2 — the onboarding checklist can show a live sync hint. GET /api/shopify/onboarding-status gained an additive catalogSync field (resolveProductSyncState(latestSyncJob), the same #205 ProductSyncState), which the #238 onboarding checklist uses to show a “syncing now” / “sync failed” hint on its sync-products row instead of a static checkbox.

Onboarding checklist: five dead signals wired + N/A steps (#321)

lib/onboarding/steps.ts is the #238 checklist registry + pure read model (effectiveStepView/isSignalDone). The registry’s steps 4-8 (categories, frontback, upsell, discount, emails) were missing signalKey even though isSignalDone always had their derivations — the switch hit case undefined and returned null, so a fully-configured store still rendered those rows as unchecked. #321 declares signalKey on all five:

  • categories: resolved + overridden + unknown > 0 && unknown === 0
  • frontback: apparel > 0 && withSelection > 0
  • upsell: upsell.configured (bundle enabled)
  • discount: discount.configured (discount enabled AND an active code)
  • emails: emails.leadCaptureEnabled || emails.klaviyoConnected

Consequence: these five now auto-check with the АВТО badge, get the honest not_detected override if a hand-checked step’s config is later disabled, and can no longer be hand-checked — only appearance and billing remain click-to-check (no signal exists for them). No new persistence: auto-done stays a render-model computation off the live probe, same as the original four AC2 steps — nothing is written to stores.onboarding_state for a signal-done step.

New StepView value "not_applicable" + pure isStepApplicable(step, signals). Today only frontback is conditional: a synced catalogue (syncedProducts > 0) with frontback.apparel === 0 renders N/A (muted skip-style row, dash glyph, reason hint, no CTA/skip/journey affordances). progressStats (panel-model.ts) drops N/A steps from the denominator entirely (11 → 10 on a no-apparel store). not_applicable wins over a done/skip entry and over D5 version resurfacing — there’s nothing left for the merchant to review. An empty catalogue deliberately stays applicable (no verdict before the first sync knows what the store sells), and signals === null (probe failure) never auto-N/As — the entry decides (AC13 unchanged). checklistProgress (checklist-progress.ts, the header-pill predicate) is untouched — it stays entry-only; the pill’s honest count instead flows through progressStats(computePanelViews(...)), so the N/A exclusion reaches it automatically.

Real-world effect, measured against prod: Icedout (jewellery, 545 products, zero hand-checked entries) went from 4/11 (36%) to 7/11 (64%) — discount/emails/upsell now auto-check; categories stays honestly open (7 unresolved products), and frontback stays applicable there because the FAMILY category resolver counts its two “Сет” bundle products as apparel. Proof: 05_tasks/proof/321/ (committed prod snapshot + replay test src/lib/onboarding/icedout-321.prod-replay.test.ts).

Onboarding checklist: the merchant’s word completes it + hide forever (#322)

Owner decision (2026-08-14): the merchant can hand-mark ANY step, one action marks everything remaining, and the checklist then hides everywhere, permanently. Three mechanics:

  • The merchant’s claim wins the read model. effectiveStepView rule 5 changed: a by: "click" done entry stands even when the probe signal is false — the row stays green and counted, with the honest hint “Marked done by you — not detected automatically”. Only a stale by: "signal" entry (the AC14 reinstall case — our past claim, not the merchant’s) still re-evaluates to not_detected. In the UI every open/resurfaced row’s status circle is a button (step-row.tsx), and a hand-checked (non-АВТО) ✓ is clickable to un-mark.
  • Complete & hide, one round trip. completableSteps(views) (panel-model.ts) = ALL steps in view open/resurfaced/skipped, signal-backed included. The footer action sends ONE PATCH /api/shopify/store-settings with onboardingCompleteSteps: string[] (#176 allowlisting: registry ids, deduped, unknown id = 400) + onboardingCollapsed: true. The server fans out one atomic merge_store_onboarding_step RPC per step via Promise.allSettled (never a whole-map read-modify-write — the #238 AC15 lost-update guarantee holds), stamps v/at/by:"click" itself, answers with onboardingComplete: { marked, failed }, and persists the _collapsed meta (setOnboardingCollapsed) ONLY when the sweep fully landed — hiding over a partial failure would bury the broken state; the panel’s retry re-sends exactly the failed ids.
  • Hidden means gone, everywhere. With _collapsed in stores.onboarding_state the panel renders null (no panel, no pill) and the top-nav header pill dies through the same checklistProgress predicate. The way back is the support bubble’s Setup guide (?onboarding=open), which clears the flag via onboardingCollapsed: false — decided once at first ready render, so a completion performed during a ?onboarding=open session is not immediately un-hidden. The ⌃ Hide button stays the per-browser localStorage collapse-to-pill; a NATURAL 100% (no flag) still renders the ✓ pill.

Proof: 05_tasks/proof/322/ + proof-322-complete-checklist.js (live dev harness); unit layer in steps.test.ts / panel-model.test.ts / store-settings/route.test.ts.

Chunked passes (#193)

runShopifyProductSync(shop, { afterCursor, maxPages }) can stop after N pages of 50 and return nextCursor; call it again with that cursor to continue. Both options are optional — omitted, it is the same unbounded pass it always was (nextCursor: null).

This exists because the one-shot sync cannot finish a large catalogue. Measured on prod: 507 products = 198 s, 451 = 169 s (~0.39 s/product), so the 2,822-product store needs ~18 minutes — and its last three sync_jobs rows are still stuck at status = running, which is why 2,739 of its rows never got their metadata back after the #190 wipe.

Two details that make resuming safe:

  • the products query pages by sortKey: ID (immutable), not UPDATED_AT — otherwise a product edited between chunks moves in the ordering and the resumed cursor skips it;
  • ensureShopifyWebPixel runs only on the pass that reaches the end (nextCursor === null), and the single-flight key includes the cursor so two different chunks are not collapsed into one run.

Driver: scripts/backfill/catalog-metadata-backfill.mjs (chunk loop over POST /api/admin/sync, which now sets maxDuration = 300); scripts/backfill/selftest.mjs proves the loop against a stub.