---
name: build-through
description: End-to-end delivery protocol for any software build request ("build me", "create an app", "add this feature", "implement this"). Captures the request verbatim, discovers the real environment, asks only load-bearing questions, writes a contract of record with mechanical checks, provisions real infrastructure first, runs a separate maker and critic, verifies the actual artifact, and reports every contract row. Use it whenever the user asks for software to be built, not for lookups or copy edits.
---

# Build-through: built all the way to the end

You carry a build from the requester's exact words to a verified result. You do not stop at a plan, a prototype, a local demo, a nice-looking screen, a passing unit test, a branch or a deployment claim. You finish when every contract row has a reproduced verdict and evidence.

This file is the short version. `BUILD-PROTOCOL.md` next to it (`.claude/skills/build-through/BUILD-PROTOCOL.md`) is the full protocol with schemas, the readiness control catalog and the artifact-specific checks. On any build that touches money, accounts, personal data, production or more than one service, read `BUILD-PROTOCOL.md` in full before Phase 2. Before Phase 0, read `CONTEXT.md` next to this file (in `.claude/skills/build-through/`). `CONTEXT.md` holds this operator's environments, approvers, authority defaults and caps, and names the folder for run records; when that slot is blank, use `builds/<build-name>/` at the repository root. If `CONTEXT.md` is missing, or a slot you need is blank, ask the user for it in the language they write in and name the slot, for example EN: "Add who approves production releases: …" or PL: „Dodaj, kto zatwierdza wydania na produkcję: …”. Until they answer, use the safe default for that slot (no remote writes, everything denied) and keep working on the rows that do not need it.

## Before you start / What you need

- **An AI coding agent with sub-agents** (required): Claude Code, Codex CLI or Cursor, signed in with your own account. The maker and the critic run in separate contexts. Claude Code: https://code.claude.com/docs/en/sub-agents
- **Git** (required): https://git-scm.com/downloads. The prompt record, the contract, parallel branches and rollback rely on it.
- **A credential store on your hosting and CI platform** (required): environment-variable or credential settings, or a password manager with a command line tool. Record in `CONTEXT.md` where credentials live, names only. The user places each value there; never ask for it in chat and never commit it. Background: https://12factor.net/config
- **Browser automation for the critic** (needed for web interfaces): Playwright, Apache-2.0, https://playwright.dev/docs/intro, or the Chrome DevTools MCP server, Apache-2.0, https://github.com/ChromeDevTools/chrome-devtools-mcp. Prove it with a positive control before relying on it.
- **Lighthouse** (optional, web surfaces): Apache-2.0, https://developer.chrome.com/docs/lighthouse/overview, for the Phase 8 performance baseline.

If one is missing, tell the user which one, link its official page, and keep working on the rows that do not need it.

## When it fires, and when it does not

Fires on: build, create, implement, set up, ship, deploy or launch software; add a feature, integration, automation, dashboard, page, service or data pipeline; redesign a user-facing surface; finish a half-built system.

Does not fire on: a pure explanation, a lookup, a code-reading question, or a copy edit that creates no software change.

Small builds get the same phases with fewer workers and checks. The prompt record, the contract, the separate critic and the row-by-row report never drop out.

## Phase 0: capture the request

Before planning, coding or summarising, copy the request **verbatim** into a prompt record with the date and the channel it arrived on. No paraphrase, no cleanup. Every later phase re-reads it. Every worker and reviewer gets its location. An agent that has not read it is not qualified to build against it.

If the request contains credentials, personal data or confidential material, keep the verbatim original access-controlled and give workers a redacted copy that keeps every requirement clause (see `BUILD-PROTOCOL.md` §7).

## Phase 1: discover the environment

Before asking anything, establish from the repository and connected tools: project root and boundaries, stack and canonical commands (install, build, test, lint, run), environments, hosting, database, auth, CI, credential storage (names and bindings only, never values), observability, rollback, and who approves what. Prove each tool you will rely on with a positive control: a read tool reads a known page, a browser tool opens a known URL. A tool that returned no error has not been proven.

Write what you found into a build-context record, each field marked verified, inferred, unknown or not applicable.

## Phase 2: extract, search, interrogate, size

1. **Extract every clause.** Number each distinct behaviour, flow, data need, integration, visual or device requirement, deployment expectation, constraint and non-goal. Offhand remarks count. Do not merge clauses because they sound related.
2. **Search before asking.** Conversation, repository, docs, tracker, build context, config, connected tools. At least three different retrieval attempts before a factual question is eligible.
3. **Ask once.** One batch of at most four ranked questions, closed choices where possible, each naming the fork and what it changes. Eligible: two readings produce different products; "done" cannot become a mechanical test; the choice changes security, privacy, money, architecture, public output or recurring cost; the action is hard to reverse. Everything cheap and reversible becomes a written assumption. On "use your judgment", record the assumptions and continue.
4. **Size it.** Estimate files, subsystems, provisioning surfaces, verification volume and evidence output. Pick one mode: one worker fits; split into units that each fit a fresh worker context; or run a mapping pass first when the build is too big to contract in one go.

## Phase 3: the contract of record

Write the contract. It defines "done". Row 0 is the risk tier:

- R1 static or local: no accounts, no persistent user writes, no money.
- R2 stateful or user-facing: accounts, persistent user data, privileged third-party access.
- R3 high impact: money, sensitive data, webhooks, uploads, background jobs, tenant boundaries, destructive automation.
- R4 scaled or regulated: multi-tenant scale, regulated data, formal on-call and recovery duties.

Row 0 also records the authority envelope: where ordinary delivery may happen (no remote writes / branch or review only / preview / staging / production) and, separately, whether migration, destructive data action, external communication, spend, credential change and production traffic are denied, prepare-only or authorized. With nothing set, default to no remote writes and everything denied.

Then one row per clause:

| # | Requirement (requester's words where possible) | Mechanical verification | Status |
|---|---|---|---|

A verification method returns pass or fail with evidence: a URL returns a status and body; an unauthorised request is refused and an authorised one succeeds; a real flow completes in the running environment; the canonical record exists after the action; the screen works at the contract's viewports; genuinely empty data renders the right empty state. "Looks right", "implemented" and a maker's self-report are not evidence.

Pair every refusal check with a positive control, so a denial proves the guard works rather than a broken system. Before trusting any new check, run it once on a known-good input and once on a known-bad one.

Rows may be added. Rows are never silently dropped: a row that cannot be met becomes a named blocker and stays in the report.

## Phase 4: provision real infrastructure first

Before any feature UI, provision or validate the real vertical slice in dependency order: source control, isolated environments, CI, data and migrations, auth, credential injection, hosting, DNS if in scope and authorised, observability, test identities and data, rollback. Record every resource you create with its purpose, owner, cost, and teardown step.

**No mock data in a real product.** Fixtures and seeds live in tests and disposable environments, labelled. A real surface uses real integrations, controlled test records, or an honest empty state with the blocker named.

**Prove the delivery path early.** Ship a trivial artifact through the real pipeline to an isolated target before features accumulate, if the authority envelope allows it. If it does not, build the artifact, run a dry run, and record the remote smoke deploy as a human-gated row.

**A missing credential does not stop the other rows.** Check the credential stores and bindings named in `CONTEXT.md` first. If it is not there, ask the human to place it in the credential store (never to paste it into chat), and keep working on every row that does not need it. Never invent a credential and never write one into a file, a log or the contract. Generate an application-owned random value only when `CONTEXT.md` grants that; provider-issued keys, paid accounts, legal terms, identity checks and consent screens are human gates.

## Phase 5: design preview gate (only when it applies)

Only for a genuinely new or materially redesigned user-facing surface. Produce several materially distinct visual previews of the key screen (count in `CONTEXT.md`), each different in hierarchy, layout, type and palette, not just colour. Label them, give one line of rationale each, and build no production UI until the approver picks. Keep working on every non-UI row meanwhile. A small change inside an existing design system skips this gate; record the skip.

## Phase 6: decompose and run the maker and critic loop

- Split by independent seams, sized so each unit fits a fresh worker. Independent units run in parallel; dependent ones run as a pipeline. Give parallel makers isolated branches, data and test accounts.
- Set hard caps before the first maker starts: iterations, wall time, cost, parallel workers, repeated-failure limit per row. The maker cannot raise its own caps.
- **Maker:** re-reads the prompt record and its rows, traces every caller before touching shared code, reuses what exists, makes the smallest correct change, adds the smallest runnable regression check, runs the canonical checks, reports with evidence.
- **Critic:** a separate agent in a fresh context. It audits whether each check can actually catch the defect, drives the running artifact (browser, device, CLI or API), reads back canonical state, tests positive and negative controls, checks sibling surfaces, and returns a verdict per row with evidence.
- **Fix the class, not the instance.** On a failed row, find the shared cause, every caller and every sibling surface, fix them all, and add a regression check. A fresh critic re-runs the row and its siblings.
- After integration, unit evidence no longer counts. Re-run the full suite on the exact integrated candidate.
- **Fail closed.** An unavailable checker, browser, environment or reviewer means "not exercised", never "passed". Stop a row after two materially different fixes fail, and report it.

## Phase 7: review gates

Major features, architecture changes, migrations, security boundaries, R3 and R4 builds, and production releases get a sectioned independent review: one reviewer per section (contract fidelity; architecture and data; UI and accessibility; backend and integrations; infrastructure and rollback; security and privacy; tests and evidence), run in parallel. The worst unresolved verdict controls the release. An uncovered surface is a finding.

Money, auth, tenant boundaries, sensitive data, state-mutating endpoints, webhooks, uploads and schema migrations also get an invariant audit: abuse vectors, invariants, cross-feature linkage, property tests, replay and concurrency, dependency failures, and a valid-user adversary pass. An open critical finding blocks release.

## Phase 8: readiness and live verification

Select the readiness controls that apply to the Row 0 tier (`BUILD-PROTOCOL.md` Appendix A). Every selected control cites a place: a file, a command, a number, a URL. A bare "yes" fails. Every skipped control states the architecture fact that makes it inapplicable.

Verify the artifact that actually ships, in the highest environment you are authorised to use: the deployed URL, the packaged app, the installed package. Not localhost, not a dev build. Use a real session where credentials allow, read results back from the source of truth, check empty states against genuinely empty data, check the contract's viewports, and measure performance against a baseline on any web surface. Say plainly what you could not exercise and why.

Deploy only inside the authority envelope. Anything outside it is a human gate: state it in one sentence with the exact action needed, and keep working on every independent row.

## Phase 9: completion report

Quote **every** contract row, in order, with its verdict and evidence:

```
Row 7 — "an admin can invite a teammate by email"
VERIFIED — invite sent to a controlled test address, record present in the invitations table, accept flow driven in the deployed app at both viewports. Evidence: <path>
```

```
Row 12 — "text message on every new order"
BLOCKED — no messaging provider account; creating one needs a payment method (human gate).
Unblocks with: the owner creates the account; then the credential is bound in the runtime and row 12 is re-verified.
```

- A report that omits a row is a failed report.
- "Done on a branch" is not done when the contract says deployed.
- Every incomplete row names its reason and the exact action that unblocks it.
- Write for someone who watched none of the build: outcome first, terms spelled out, one plain clause per file, URL or row.

Then update the tracker named in `CONTEXT.md` and leave a handoff record, so the next operator can continue without reconstructing the build from chat.

## Why each phase exists

- The verbatim prompt record and the contract give every agent the same definition of "done", so no requirement depends on one agent's memory.
- Provisioning before features means every screen runs on the real backend, and every check grades the real product.
- A separate critic on the running artifact catches what the author's own framing hides, including the same defect on sibling screens.
- An explicit authority envelope lets the agent finish everything it is allowed to do and stop cleanly at everything it is not.
