---
name: correction-log
description: Makes any recurring agent workflow (a report, an edit, a deploy, a content pipeline) learn from its corrections through an append-only log the agent reads before every run and must grow before the run can close. Use it when the same workflow runs again and again and a human keeps correcting the same kinds of mistakes.
---

# Correction log: a workflow that learns from its mistakes

A recurring workflow repeats its mistakes unless something carries the lesson from one run to
the next. This skill gives the workflow that memory: one append-only log per workflow, read end to
end before every run, written to before every run closes, and backed by gates that measure instead
of assert.

Read `CONTEXT.md` next to this file first. It names the workflow, where its log lives, who counts
as a correction source, the gates, and who does the independent re-check at hand-off. If
`CONTEXT.md` is missing or a field you need is blank, ask the user once, in the language they write
in, then write the answer into `CONTEXT.md` so nobody asks again. For example:

- English: "Add your workflow's stages and the gate that ends each one."
- Polish: „Dodaj etapy swojego procesu i bramkę, która kończy każdy z nich."

The template ships with no workflow, thresholds or owner. Never invent them.

## Before you start / What you need

- **An AI coding agent** that reads a markdown instruction file and can run shell commands: Claude
  Code, Codex CLI or Cursor.
- **A bash shell with grep** to run `log-gate.sh`. macOS and Linux have both. On Windows, use WSL.
- **A second agent session** for the hand-off check: a new session of the same agent is enough, as
  long as it did not build the output.
- **A filled-in `CONTEXT.md`** and a `CORRECTION-LOG.md` copied from the template.

## The four files

- `SKILL.md` (this file): the method. It changes rarely.
- `CONTEXT.md`: the operator's details. It changes when the workflow changes.
- `CORRECTION-LOG.md`: the workflow's memory. It only ever grows. Start from
  `CORRECTION-LOG.template.md` and delete its two example rows.
- `log-gate.sh`: a tiny script that refuses to close a run unless the log grew.

One log per workflow. A weekly report and a deploy pipeline get separate logs, because a rule
learned on one rarely applies to the other and a mixed log gets skimmed instead of read.

## Before the run: read the whole log

1. Open the workflow's `CORRECTION-LOG.md` and read every row, top to bottom. Do not search it
   for keywords and do not read only the latest rows. The log holds mistakes already paid for;
   the point of the read is to avoid paying for one of them twice.
2. Record the baseline so the close check can prove the log grew:

   ```bash
   ./log-gate.sh begin path/to/CORRECTION-LOG.md
   ```

3. If `CONTEXT.md` lists a run ledger, read its last three rows. They set the time you expect
   each stage to take on this run.
4. Before you start, state in one line which log rules apply to this run. A rule you cannot
   connect to this run still gets checked in the hand-off log re-read below; the line only proves the
   read happened.

## During the run: gates measure, they do not assert

Each stage of the workflow ends in a gate: a check that measures the thing that went wrong
before. A stage is finished when its gate passes with numbers written down, not when the output
looks done.

Every gate follows three rules:

1. **Write numbers or nothing.** A gate records what it measured (row counts, totals, durations,
   diffs, error counts) next to its verdict. A bare "PASS" is not evidence. If the gate cannot
   produce the number, it does not produce a verdict either.
2. **"Could not measure" is its own exit state.** Use three outcomes: pass, fail, and could not
   measure. Could not measure halts the stage just like a fail; it is never rounded up to a pass.
   A missing input, a timeout, or a tool that returned nothing all land here.
3. **Watch it fail before you trust it.** Before a gate goes into service, and after any change
   to it, run it against a deliberately broken copy of a real output: a total that no longer
   adds up, a missing section, a truncated file, whatever the gate claims to catch. The gate must
   fail on that copy. A gate that passes a broken copy is decoration. Write the broken-copy test
   down in `CONTEXT.md` so it can be repeated.

If a gate fails twice on the same class of problem in one run, stop. Report what the gate
measured, which numbers sit out of tolerance, the options, and the one you recommend. Do not
lower a threshold to reach green; a threshold moves only when the owner named in `CONTEXT.md`
says so, and that change is itself a log row.

## Hand-off: separation of duty

Nothing reaches the human until the hand-off check passes and its results are written to a
hand-off note next to the output.

**The builder** (the agent that produced the output) runs every gate in `CONTEXT.md` plus:

- **Log re-read.** Walk the correction log against the finished output, row by row. Each
  past rule is checked against this output, not assumed to still hold.

**The checker** (a second agent that did not build the output, named in `CONTEXT.md`) re-runs
the cheap decisive checks itself before the human sees anything. Cheap means seconds, not
minutes. Decisive means a failure there would make the output wrong, not merely less polished.
`CONTEXT.md` lists which checks those are.

The builder's report does not count as the check. "The agent said it passed" is a claim, not a
check. If the checker cannot run a check, the output does not ship; the hand-off note says which
check could not run and why.

## After the run: write corrections back before closing

The run does not close until the log grew. Append one row per lesson, in this shape:

```
| Date | Source | Correction |
| YYYY-MM-DD | user / QA / post-mortem | Symptom: ... Cause: ... Rule: ... |
```

- **Source** says where the correction came from: `user` (the human corrected the output), `QA`
  (a gate or the checker caught it), or `post-mortem` (found after delivery).
- **Symptom** is what was visibly wrong. **Cause** is why it happened. **Rule** is the
  instruction that now applies to every future run, written as an imperative a new agent could
  follow with no other context.

What counts as a row:

- every mistake the human corrected, however small;
- every workaround you had to invent;
- every threshold the owner moved;
- every gate that could not measure;
- every gate that let a broken copy through.

Then close the run:

```bash
./log-gate.sh close path/to/CORRECTION-LOG.md
```

It exits non-zero unless the log has more rows than at `begin`. A run with genuinely nothing to
learn still adds a row: `Symptom: none. Rule: no change.` That row keeps the close honest and
records that the run happened.

## Fix the class, not the instance

When you write the rule column, name the class of problem, not the single case. "Last month's total
was wrong" fixes one report. "Totals are recomputed from line items and compared with the
source's own total before the report is written" fixes every future report. If you can only
write the instance, the cause column is not finished yet.

## A correction given twice is a pipeline failure

Before you append a row, search the log for the same class. If an earlier row already covers it,
the log did not work for that lesson, and another row will not work either. Instead:

1. Turn the rule into a gate: a check that measures it on every run, with a broken-copy test.
2. Or move it into the body of the skill or workflow instructions, where it is followed rather
   than remembered.
3. Then append a row that says so: `Rule: repeat of <earlier date>; now enforced by <gate or
   section>.`

A log where the same class appears three times is a sign the escalation step is being skipped.

## Optional: the run ledger

If `CONTEXT.md` enables it, append one row per run to a separate `RUN-LEDGER.md`:

```
| Date | Run | Stage timings | Log rules applied | Rows added | Note |
```

Record wall time per stage from the start of the stage to its gate passing. Write `no data` for
anything you did not measure; never estimate. Read the last three rows before each run. A stage
that runs clearly slower than the previous comparable run is a finding: name the cause, and if it
repeats, it becomes a correction-log row. Over time the ledger shows whether the workflow is
getting faster and cleaner, or only growing a longer log.

## Maintaining the log

- **Append only.** Never edit or delete a past row. If a rule turns out wrong, append a row that
  supersedes it and names the earlier date.
- **Keep rows short.** One to three sentences. The log is read end to end every run, so padding
  costs reading time on every run.
- **Graduate rules.** When a rule has become a gate or a section of the instructions, leave the
  row in place and note the promotion in a new row. The history stays; the enforcement moves.
- **Split when a log grows too long to read in one sitting.** Move rules that are now enforced by
  gates into a short "enforced elsewhere" list at the top, and keep the table for live lessons.

## Failure modes to watch for

| What you see | What it means | What to do |
|---|---|---|
| The run closes without a new row | The close check was skipped | Run `log-gate.sh close` as the last step, every run |
| Rows describe one case each | The class was never named | Rewrite the rule column as a class before closing |
| The same class appears twice | The log is not changing behaviour | Promote the rule to a gate or the instructions |
| A gate only ever prints PASS | It records no numbers or was never tested | Add the numbers and run the broken-copy test |
| The checker always agrees with the builder | The checker is reading the builder's report | Make the checker re-run the checks from the raw output |
| Thresholds drift over time | Someone lowered a bar to get green | Only the named owner moves a threshold, and it is logged |

## When not to use this

- A one-off task that will not run again. There is no next run to learn.
- A workflow with no human or gate that can say what went wrong. Add a gate first.
- Personal taste that changes every run. Log only corrections that should still hold next time.
