---
name: adversarial-review
description: Runs an independent, adversarial review gate before any major artifact, risky code change or production deploy is locked, and adds a seven-section invariant audit for money paths, auth, tenant boundaries, state mutation and schema migrations. Use it when the user says "ship it", "ready for review", "what could go wrong" or "before we merge", or when a change touches billing, login, permissions or shared data.
---

# Adversarial review gate

The agent that built something is the worst judge of it. It reads its own work with the same frame it
wrote it with, so the gaps it missed while building stay invisible while reviewing. This skill puts a
separate reviewer between "done" and "locked", gives that reviewer a hostile brief, and lets its verdict
stop the work.

Read `CONTEXT.md` next to this file before the first review. It names your reviewer, where reviews are
saved, and which parts of your product count as money, auth and tenant surfaces.

If `CONTEXT.md` is missing, or a field you need is blank, stop and ask the user for it in the user's
language before the first dispatch. Ask for one field at a time and say what it is for, for example:
EN "Add your money surfaces: which files, services or endpoints move or record money in your product?",
PL „Dodaj swoje miejsca, w których płyną pieniądze: które pliki, usługi albo endpointy liczą lub zapisują płatności?”.
Never guess a surface list from the code alone; confirm it with the user and write it into `CONTEXT.md`.

## Before you start / What you need

- **An AI coding agent that runs subagents in parallel** (required): Claude Code
  (https://code.claude.com/docs/en/sub-agents), Codex CLI (https://developers.openai.com/codex/subagents)
  or Cursor (https://cursor.com/docs/agent/subagents). Give each reviewer read, search and run-tests tools
  only: in Claude Code and Cursor through the subagent definition, in Codex CLI with `--sandbox read-only`.
- **A second model as reviewer** (optional, recommended): any command-line tool from another vendor.
  One public example is Codex CLI (https://developers.openai.com/codex/cli, source
  https://github.com/openai/codex, Apache-2.0): install with `npm install -g @openai/codex`, run `codex`
  once and sign in through its own login, then review with
  `codex exec --sandbox read-only "<reviewer prompt>"`. Keep login details in the tool's own store,
  never in the chat or the repository.
- **Node.js with npm** (optional, https://nodejs.org/en/download): only to install Codex CLI or fast-check.
- **Python with pip** (optional, https://www.python.org/downloads/): only to install Hypothesis.
- **Git** (required, https://git-scm.com/downloads): `git diff --stat <base>...HEAD` and
  `git diff -U0 <base>...HEAD` give the file list and line ranges for step 1.
- **A property-testing library** (only for the invariant audit): Hypothesis for Python
  (https://hypothesis.readthedocs.io, `pip install hypothesis`) or fast-check for JavaScript and
  TypeScript (https://fast-check.dev, `npm install --save-dev fast-check`).

## Rule zero: the maker is never the checker

- The reviewer did not write the change. A second pass by the same agent in the same conversation is
  not a review.
- Prefer a reviewer from a different model or tool, where available (see `CONTEXT.md`). A different
  model family misses different things. If you only have one model, use a fresh session with no memory
  of the build and a read-only tool set.
- The reviewer has no write access. Give it read, search and run-tests tools only. A reviewer that can
  edit will fix what it finds, and then nobody has reviewed the fix.
- The reviewer stops after the verdict. Fixing is the maker's job.

## When the gate fires

Dispatch a review, without being asked, when any of these happen:

1. A decision record or architecture artifact is saved or locked: a design decision, a system map, a
   schema, an API or action inventory, or any document moving from research into architecture.
2. A major rewrite of a locked rule set: a decision catalog rewrite, a revised decision list, a schema
   migration plan.
3. A money-path code change: ledger, balances, charges, refunds, billing webhooks, usage metering.
4. An auth or security boundary change: tenant isolation, token verification, identity-provider
   integration, row-level access policies, session handling.
5. A milestone ships: a first end-to-end version, or a release that adds a major capability.
6. A production deploy: any push that changes a customer-facing surface.

Also fire when the user asks for sign-off in any words, for example "ship it", "ready for review",
"what could go wrong" or "before we merge", or refers to merging or finalising a major artifact.
Add your own phrases to the "Trigger phrases" field in `CONTEXT.md`.

## When the gate does not fire

- Pure documentation fixes: typos, formatting, broken links.
- Notes, status logs, meeting summaries.
- Planning files explicitly marked as drafts and not yet locked.
- Conversation that produces no committed artifact.
- Routine script runs that change nothing structural.

## Step 1: freeze the change

List exactly what changed: file paths and line ranges for every new or modified block. No list, no
dispatch. A reviewer told to "check the project" skims everything and returns generalities.

## Step 2: section the review

One reviewer per section, all dispatched in parallel in a single message. Never one reviewer for
multi-part work: a lone reviewer thin-spreads across the whole surface and misses what a bounded
reviewer catches.

- Cut by domain or concern: frontend, API, billing, auth, webhooks, migrations, data integrity.
- For user-facing work, cut by user journey instead: the path a real person walks end to end.
- Name which section owns each seam between sections. The seams are where coverage goes missing.
- Write the sectioning plan down before dispatch: "Section A covers X, section B covers Y, the seam
  between A and B belongs to A."

## Step 3: dispatch with the full contract

Every reviewer prompt contains all seven parts. Leave one out and the review is not valid.

1. **What changed**: exact file paths and line ranges for this section.
2. **Locked context**: every locked decision the artifact must respect (architecture decisions, infra
   choices, the auth design, known facts about the data).
3. **Source-truth files**: the research, specs or extraction documents to cross-check fidelity against.
4. **Adversarial charter**: "Kill weak rules. Find missing constraints. Flag wishful enforcement (a rule
   written down with nothing that checks it). Find stack incompatibilities. Point at every assumption
   the author did not verify. Don't be polite. Don't praise."
5. **Output path**: save the review to the review folder from `CONTEXT.md` as
   `review-<date>-<topic>-<section>.md`, and reply with the path plus a short summary.
6. **Structured verdict**: exactly one of `BORINGLY RELIABLE`, `CRITICAL GAPS`, `MAJOR REWRITE`.
7. **Sectioning plan**: how the work was split, which section this reviewer owns, and confirmation that
   each section has exactly one reviewer.

Prompt skeleton:

```
Role: adversarial reviewer. You did not build this. You are not polite.

What changed: [paths + line ranges]
Locked context: [decisions this must respect]
Source of truth: [files to cross-check]
Your section: [boundary; review only this]
Sectioning plan: [all sections and who owns each seam]

Charter: kill weak rules, find missing constraints, flag wishful enforcement,
find stack incompatibilities, name every unverified assumption. No praise.

Return:
- Coverage map: what you actually inspected
- Findings: severity (CRITICAL / MAJOR / MINOR) + file:line each
- Verdict: BORINGLY RELIABLE | CRITICAL GAPS | MAJOR REWRITE
- If BORINGLY RELIABLE: three specific things you tried to break and could not.
  Without these three, the verdict is void.

Save to: [review folder]/review-[date]-[topic]-[section].md
Reply: path + short summary.
Stop after the verdict. Do not fix anything.
```

## Step 4: synthesise

- The combined verdict is the **worst** section verdict. One `CRITICAL GAPS` anywhere blocks the lock,
  even if every other section is clean.
- Merge the coverage maps. Any surface no section covered is itself a finding: dispatch a reviewer for it.
- Reconcile conflicting findings. Never drop one silently.
- Reject a `BORINGLY RELIABLE` that lists nothing the reviewer tried to break. A polite review is theatre.

## Step 5: act on the verdict

- `CRITICAL GAPS` or `MAJOR REWRITE`: do not proceed past the gate. Fix each finding concretely, then
  re-dispatch a reviewer on the fixed section only. Repeat until that section returns `BORINGLY RELIABLE`.
- `BORINGLY RELIABLE` on every section: proceed to lock, merge or deploy. Keep the review files as the
  audit trail and cite them from later decisions.
- Before fixing, open every cited `file:line` and confirm the finding yourself. Reviewers hallucinate
  findings; "fixing" one breaks working code.
- Three fix-and-review rounds on one section mean the design is wrong, not the code. Step back to design.

## Invariant audit: money, auth, tenants, state, migrations

The review asks "does this hang together". The invariant audit asks "under hostile input, do the
invariants hold". Run both before locking any change in these areas (your product's concrete surfaces
are listed in `CONTEXT.md`):

- **Money path**: credit ledgers, balances, idempotency keys, charge, refund and dispute flows, billing
  webhooks, spending limits, plan or permission-level changes.
- **Auth and session boundary**: token verification, identity-provider webhooks, session refresh and
  rotation, multi-factor auth, password reset, OAuth callbacks, support impersonation, API credential
  issuance and revocation.
- **Tenant boundary**: new admin actions, cross-tenant features, tenant deletion and right-to-erasure,
  subscription-cancellation cascades, workspace switching.
- **State mutation**: any new create, update or delete endpoint on tenant-scoped data, background jobs
  that mutate billing or tenant state, inbound webhook handlers, file uploads, rate-limit changes.
- **Schema migrations**: any migration touching tenant-scoped or money-bearing tables, new tenant-scoped
  tables, index changes on hot tables.

It does not fire on documentation, formatting, non-auth styling, notes or status logs.

Fill `INVARIANT-AUDIT.template.md` (next to this file). All seven sections are mandatory:

1. **Abuse vectors**: "if an actor with this access does this action, they achieve this outcome". At least five per money-path
   or auth change.
2. **Invariant locks**: properties that must hold for every tenant under every concurrent invocation
   (balance never negative, one idempotency key produces one ledger entry, tenant A never reads tenant
   B, a failed action never charges).
3. **Cross-feature linkage**: for each invariant, every other feature that touches the same state, and
   whether this change breaks its assumptions.
4. **Property-based tests**: concrete properties, not examples, written for a property-testing library
   such as Hypothesis (Python) or fast-check (TypeScript/JavaScript).
5. **Replay, concurrency and races**: duplicate webhook delivery, many concurrent requests on one
   idempotency key, a job crashing mid-pipeline and then retrying.
6. **Failure modes**: database drop mid-transaction, job timeout, provider error, cache eviction before
   retry, unreachable key endpoint, a rotated webhook signing credential.
7. **Counterfactual adversary pass**: "as a valid user on plan X, what is my highest-leverage attack
   right now?" Each attack names the target invariant, entry point, chain of steps, and a status:
   `blocked`, `partial` or `open`.

**Any vector marked `open` or `partial` blocks the lock.** Build the mitigation, re-run the audit on it,
and lock only when every vector is `blocked` or accepted with the residual risk written down and signed
off by a human. Save the audit to the audit folder from `CONTEXT.md` as
`invariant-audit-<date>-<feature>.md` and reference it from the lock.

Why this exists: a design review can pass while a double charge, a cross-tenant read or a replayed
webhook still slips through. Each of those is a top-severity incident.

## Failure modes of the gate itself

- **Same-agent review.** The builder "reviews" its own work in the same thread. Fix: a separate reviewer,
  always.
- **One reviewer for everything.** Fix: section, and treat any uncovered surface as a finding.
- **Reviewer with write tools.** It repairs what it found and the repair goes unreviewed. Fix: read-only.
- **Vague brief.** No file list, no charter, no verdict format. Fix: the seven-part contract, every time.
- **Trusting an unchecked finding.** Fix: open the `file:line` before changing anything.
- **Locking past a gap** "because the other sections were fine". Fix: worst verdict wins.
