Correction log: an AI agent that learns from its mistakes

Your recurring workflow remembers every correction and will not close a run until it writes the correction down.

Operations, intermediate. Published

What it does

Before each run of a workflow, the agent reads the whole correction log, from the first row to the last. Every stage ends in a gate that records measured numbers, and "could not measure" halts the stage just like a failure. Before the output reaches you, a second agent that did not build it re-runs the cheap decisive checks itself. At the end, the agent appends every correction to the log: the symptom, the cause, and the rule for next time. The log-gate.sh script refuses to close the run unless the log grew. When the same correction shows up a second time, the agent moves it into a gate or into the workflow instructions instead of adding another row.

When to use it

When to use it

  • An agent runs the same workflow regularly: a report, a publish step, a deploy, an edit, a data export.
  • You correct the agent in the same places run after run.
  • You want the output checked before you see it, not after.

When not to use it

  • A one-off task that will not run again.
  • A workflow where no person and no check can say what went wrong. Add a gate first.
  • Preferences you change on every run.

Decision table

SituationWhat the skill does
You correct an output for the first timeThe agent appends a row: symptom, cause, and a rule for the whole class of problem
The same correction comes backThe agent moves the rule into a gate or the instructions and logs the move
A gate cannot measure somethingThe stage halts, and the missing measurement becomes its own row
The agent says everything passedA second agent re-runs the cheap decisive checks before hand-off
A run takes longer than the last oneThe run ledger shows the slower stage, and the agent names the cause

Template

The method: read the log before every run, gates that record numbers, a second check at hand-off, corrections written back before the run closes.

SKILL.mdDownloadSKILL.md
---
name: correction-log
description: Makes any recurring agent workflow (a report, an edit, a deploy, a content pipeline) learn from its corrections through an append-only log the agent reads before every run and must grow before the run can close. Use it when the same workflow runs again and again and a human keeps correcting the same kinds of mistakes.
---

# Correction log: a workflow that learns from its mistakes

A recurring workflow repeats its mistakes unless something carries the lesson from one run to
the next. This skill gives the workflow that memory: one append-only log per workflow, read end to
end before every run, written to before every run closes, and backed by gates that measure instead
of assert.

Read `CONTEXT.md` next to this file first. It names the workflow, where its log lives, who counts
as a correction source, the gates, and who does the independent re-check at hand-off. If
`CONTEXT.md` is missing or a field you need is blank, ask the user once, in the language they write
in, then write the answer into `CONTEXT.md` so nobody asks again. For example:

- English: "Add your workflow's stages and the gate that ends each one."
- Polish: „Dodaj etapy swojego procesu i bramkę, która kończy każdy z nich."

The template ships with no workflow, thresholds or owner. Never invent them.

## Before you start / What you need

- **An AI coding agent** that reads a markdown instruction file and can run shell commands: Claude
  Code, Codex CLI or Cursor.
- **A bash shell with grep** to run `log-gate.sh`. macOS and Linux have both. On Windows, use WSL.
- **A second agent session** for the hand-off check: a new session of the same agent is enough, as
  long as it did not build the output.
- **A filled-in `CONTEXT.md`** and a `CORRECTION-LOG.md` copied from the template.

## The four files

- `SKILL.md` (this file): the method. It changes rarely.
- `CONTEXT.md`: the operator's details. It changes when the workflow changes.
- `CORRECTION-LOG.md`: the workflow's memory. It only ever grows. Start from
  `CORRECTION-LOG.template.md` and delete its two example rows.
- `log-gate.sh`: a tiny script that refuses to close a run unless the log grew.

One log per workflow. A weekly report and a deploy pipeline get separate logs, because a rule
learned on one rarely applies to the other and a mixed log gets skimmed instead of read.

## Before the run: read the whole log

1. Open the workflow's `CORRECTION-LOG.md` and read every row, top to bottom. Do not search it
   for keywords and do not read only the latest rows. The log holds mistakes already paid for;
   the point of the read is to avoid paying for one of them twice.
2. Record the baseline so the close check can prove the log grew:

   ```bash
   ./log-gate.sh begin path/to/CORRECTION-LOG.md
   ```

3. If `CONTEXT.md` lists a run ledger, read its last three rows. They set the time you expect
   each stage to take on this run.
4. Before you start, state in one line which log rules apply to this run. A rule you cannot
   connect to this run still gets checked in the hand-off log re-read below; the line only proves the
   read happened.

## During the run: gates measure, they do not assert

Each stage of the workflow ends in a gate: a check that measures the thing that went wrong
before. A stage is finished when its gate passes with numbers written down, not when the output
looks done.

Every gate follows three rules:

1. **Write numbers or nothing.** A gate records what it measured (row counts, totals, durations,
   diffs, error counts) next to its verdict. A bare "PASS" is not evidence. If the gate cannot
   produce the number, it does not produce a verdict either.
2. **"Could not measure" is its own exit state.** Use three outcomes: pass, fail, and could not
   measure. Could not measure halts the stage just like a fail; it is never rounded up to a pass.
   A missing input, a timeout, or a tool that returned nothing all land here.
3. **Watch it fail before you trust it.** Before a gate goes into service, and after any change
   to it, run it against a deliberately broken copy of a real output: a total that no longer
   adds up, a missing section, a truncated file, whatever the gate claims to catch. The gate must
   fail on that copy. A gate that passes a broken copy is decoration. Write the broken-copy test
   down in `CONTEXT.md` so it can be repeated.

If a gate fails twice on the same class of problem in one run, stop. Report what the gate
measured, which numbers sit out of tolerance, the options, and the one you recommend. Do not
lower a threshold to reach green; a threshold moves only when the owner named in `CONTEXT.md`
says so, and that change is itself a log row.

## Hand-off: separation of duty

Nothing reaches the human until the hand-off check passes and its results are written to a
hand-off note next to the output.

**The builder** (the agent that produced the output) runs every gate in `CONTEXT.md` plus:

- **Log re-read.** Walk the correction log against the finished output, row by row. Each
  past rule is checked against this output, not assumed to still hold.

**The checker** (a second agent that did not build the output, named in `CONTEXT.md`) re-runs
the cheap decisive checks itself before the human sees anything. Cheap means seconds, not
minutes. Decisive means a failure there would make the output wrong, not merely less polished.
`CONTEXT.md` lists which checks those are.

The builder's report does not count as the check. "The agent said it passed" is a claim, not a
check. If the checker cannot run a check, the output does not ship; the hand-off note says which
check could not run and why.

## After the run: write corrections back before closing

The run does not close until the log grew. Append one row per lesson, in this shape:

```
| Date | Source | Correction |
| YYYY-MM-DD | user / QA / post-mortem | Symptom: ... Cause: ... Rule: ... |
```

- **Source** says where the correction came from: `user` (the human corrected the output), `QA`
  (a gate or the checker caught it), or `post-mortem` (found after delivery).
- **Symptom** is what was visibly wrong. **Cause** is why it happened. **Rule** is the
  instruction that now applies to every future run, written as an imperative a new agent could
  follow with no other context.

What counts as a row:

- every mistake the human corrected, however small;
- every workaround you had to invent;
- every threshold the owner moved;
- every gate that could not measure;
- every gate that let a broken copy through.

Then close the run:

```bash
./log-gate.sh close path/to/CORRECTION-LOG.md
```

It exits non-zero unless the log has more rows than at `begin`. A run with genuinely nothing to
learn still adds a row: `Symptom: none. Rule: no change.` That row keeps the close honest and
records that the run happened.

## Fix the class, not the instance

When you write the rule column, name the class of problem, not the single case. "Last month's total
was wrong" fixes one report. "Totals are recomputed from line items and compared with the
source's own total before the report is written" fixes every future report. If you can only
write the instance, the cause column is not finished yet.

## A correction given twice is a pipeline failure

Before you append a row, search the log for the same class. If an earlier row already covers it,
the log did not work for that lesson, and another row will not work either. Instead:

1. Turn the rule into a gate: a check that measures it on every run, with a broken-copy test.
2. Or move it into the body of the skill or workflow instructions, where it is followed rather
   than remembered.
3. Then append a row that says so: `Rule: repeat of <earlier date>; now enforced by <gate or
   section>.`

A log where the same class appears three times is a sign the escalation step is being skipped.

## Optional: the run ledger

If `CONTEXT.md` enables it, append one row per run to a separate `RUN-LEDGER.md`:

```
| Date | Run | Stage timings | Log rules applied | Rows added | Note |
```

Record wall time per stage from the start of the stage to its gate passing. Write `no data` for
anything you did not measure; never estimate. Read the last three rows before each run. A stage
that runs clearly slower than the previous comparable run is a finding: name the cause, and if it
repeats, it becomes a correction-log row. Over time the ledger shows whether the workflow is
getting faster and cleaner, or only growing a longer log.

## Maintaining the log

- **Append only.** Never edit or delete a past row. If a rule turns out wrong, append a row that
  supersedes it and names the earlier date.
- **Keep rows short.** One to three sentences. The log is read end to end every run, so padding
  costs reading time on every run.
- **Graduate rules.** When a rule has become a gate or a section of the instructions, leave the
  row in place and note the promotion in a new row. The history stays; the enforcement moves.
- **Split when a log grows too long to read in one sitting.** Move rules that are now enforced by
  gates into a short "enforced elsewhere" list at the top, and keep the table for live lessons.

## Failure modes to watch for

| What you see | What it means | What to do |
|---|---|---|
| The run closes without a new row | The close check was skipped | Run `log-gate.sh close` as the last step, every run |
| Rows describe one case each | The class was never named | Rewrite the rule column as a class before closing |
| The same class appears twice | The log is not changing behaviour | Promote the rule to a gate or the instructions |
| A gate only ever prints PASS | It records no numbers or was never tested | Add the numbers and run the broken-copy test |
| The checker always agrees with the builder | The checker is reading the builder's report | Make the checker re-run the checks from the raw output |
| Thresholds drift over time | Someone lowered a bar to get green | Only the named owner moves a threshold, and it is logged |

## When not to use this

- A one-off task that will not run again. There is no next run to learn.
- A workflow with no human or gate that can say what went wrong. Add a gate first.
- Personal taste that changes every run. Log only corrections that should still hold next time.

Other files

CORRECTION-LOG.template.mdAn empty correction log with its header and two example rows to delete.
CORRECTION-LOG.template.mdDownloadCORRECTION-LOG.template.md
# Correction log: [workflow name]

Append-only. Read every row before each run; append before the run closes.
A correction given twice is a pipeline failure: promote it to a gate or the instructions.

Each row: date · source (user / QA / post-mortem) · symptom, cause, and the rule that now applies.
Never edit or delete a past row; supersede it with a new one that names the earlier date.

## Enforced elsewhere

(Rules that have become gates or instruction sections. One line each, with the gate or section name.)

## Corrections

| Date | Source | Correction |
|---|---|---|
| 2000-01-01 | user | Example row, delete it. Symptom: the weekly sales report listed refunded orders as revenue. Cause: the query read gross order totals and ignored the refund table. Rule: revenue is net of refunds; the report gate compares the net total with the payment provider's payout total and fails on any gap. |
| 2000-01-08 | QA | Example row, delete it. Symptom: the report shipped with an empty chart. Cause: the export step timed out and the chart step read a zero-byte file as "no sales". Rule: a zero-byte or missing input is "could not measure" and halts the run; it is never read as zero. |
log-gate.shA small bash script: begin records the row count, close fails unless the log grew.
log-gate.shDownloadlog-gate.sh
#!/usr/bin/env bash
# log-gate.sh: refuse to close a run unless its correction log grew.
#   log-gate.sh begin <log>   record the current row count in <log>.baseline
#   log-gate.sh close <log>   exit 0 only if the log now has more rows than the baseline
# A row is a table line that starts with a date: "| YYYY-MM-DD |".
# Exit: 0 ok · 1 log did not grow · 2 could not measure (missing log or baseline) · 3 usage.
set -u

count_rows() { grep -cE '^\|[[:space:]]*[0-9]{4}-[0-9]{2}-[0-9]{2}[[:space:]]*\|' "$1"; }

[ $# -eq 2 ] || { echo "usage: log-gate.sh begin|close <log>" >&2; exit 3; }
cmd="$1"; log="$2"; base="$log.baseline"
[ -f "$log" ] || { echo "could not measure: log not found: $log" >&2; exit 2; }
now="$(count_rows "$log")"

case "$cmd" in
  begin)
    echo "$now" > "$base" || { echo "could not measure: cannot write $base" >&2; exit 2; }
    echo "begin: $now rows recorded in $base"
    ;;
  close)
    [ -f "$base" ] || { echo "could not measure: no baseline, run 'begin' first" >&2; exit 2; }
    was="$(cat "$base")"
    case "$was" in ''|*[!0-9]*) echo "could not measure: bad baseline in $base" >&2; exit 2;; esac
    if [ "$now" -gt "$was" ]; then
      echo "close: log grew from $was to $now rows"
      rm -f "$base"
    else
      echo "close refused: log has $now rows, baseline $was. Append this run's corrections first." >&2
      exit 1
    fi
    ;;
  *) echo "usage: log-gate.sh begin|close <log>" >&2; exit 3 ;;
esac
CONTEXT.template.mdContext template: your workflow, stages, gates, thresholds, the checker and the run ledger.
CONTEXT.template.mdDownloadCONTEXT.template.md
# Correction log: context

Copy this file to CONTEXT.md next to SKILL.md and fill in every blank (______). Example lines only show the shape of an answer; delete them once you have written your own. Never put passwords, access keys or card numbers here.

## The workflow

(What recurring workflow this log serves, how often it runs, and what it delivers.)
Your answer: ______
Example (replace it, this is not your data): Weekly sales report, at the start of each week, a one-page summary sent to the shop owner.

## Stages

(The stages in order, each with the gate that ends it. One line per stage.)
Your answer: ______
Example (replace it, this is not your data): 1. Pull orders and refunds. Gate: row counts match the source export.
Example (replace it, this is not your data): 2. Compute totals. Gate: net total matches the payment provider's payout within the tolerance below.
Example (replace it, this is not your data): 3. Write the report. Gate: every section present, no placeholder text.

## Where the files live

(Path to this workflow's CORRECTION-LOG.md, to log-gate.sh, and to the run ledger if you use one.)
Your answer: ______
Example (replace it, this is not your data): reports/weekly-sales/CORRECTION-LOG.md

## Correction sources

(Who or what may add a row: the people whose corrections count, the gates that count as QA, and when a post-mortem happens.)
Your answer: ______
Example (replace it, this is not your data): user = the shop owner; QA = any gate failure or checker finding; post-mortem = any error found after the report was sent.

## Thresholds and their owner

(Each gate threshold, its value, how you derived it, and the one person allowed to change it.)
Your answer: ______
Example (replace it, this is not your data): payout match tolerance: the smallest gap your payment provider's rounding produces in a normal week. Owner: the shop owner.

## Broken-copy tests

(For each gate, the deliberately broken copy it must fail on, and when you last watched it fail.)
Your answer: ______
Example (replace it, this is not your data): totals gate: copy last week's report, delete one refund row, confirm the gate fails.

## Hand-off

(Who the checker is, which cheap decisive checks the checker re-runs, and where the hand-off note is written.)
Your answer: ______
Example (replace it, this is not your data): checker = a second agent session that did not build the report; it re-runs the totals gate and the section check from the raw export.

## Run ledger

(On or off. If on, the stages you time.)
Your answer: ______
Example (replace it, this is not your data): on; time each of the three stages above.

## Notes for the agent

(Anything else a new agent needs before its first run: tone of the output, who receives it, known quirks of the data.)
Your answer: ______

What you need

  • An AI coding agent: Claude Code, Codex CLI or Cursor (opens in new tab)

    Required · App

    Runs the workflow, reads the log and runs log-gate.sh. It needs to read a markdown instruction file and run shell commands.

    1. Install one agent from its official docs: Claude Code (code.claude.com/docs), Codex CLI (developers.openai.com/codex) or Cursor (cursor.com/docs).
    2. Place the template folder where your agent loads skills or instructions, as in the install steps.
    3. For the hand-off check, open a second, new session of the same agent. It must not be the session that built the output.
  • Bash shell with grep (opens in new tab)

    Required · Runtime · License GPL-3.0-or-later

    log-gate.sh is a bash script that counts log rows with grep.

    1. macOS and Linux ship with bash and grep. Check with: bash --version.
    2. Windows: install WSL following Microsoft's guide (learn.microsoft.com/windows/wsl/install) and run the script inside WSL.
    3. Make the script executable: chmod +x log-gate.sh.

Install

  1. Claude Code: copy the folder to .claude/skills/correction-log/ in your project, or to ~/.claude/skills/correction-log/ for every project. Claude loads it when the description matches the task, or when you type /correction-log.
  2. Codex: paste SKILL.md into AGENTS.md at the repository root or in ~/.codex/AGENTS.md.
  3. Cursor: add SKILL.md as a project rule or paste it into AGENTS.md.
  4. Any other agent: paste SKILL.md into the project instruction file or system prompt.
  5. Copy CONTEXT.template.md to CONTEXT.md next to SKILL.md and fill in every field.
  6. Copy CORRECTION-LOG.template.md next to your workflow as CORRECTION-LOG.md, delete the two example rows, and make log-gate.sh executable (chmod +x).

It's working if

  • The agent's first reply in a run lists the log rules that apply to this run.
  • A .baseline file appears next to the log when a run starts and disappears after a clean close.
  • log-gate.sh close exits with code 1 when you try to close a run without a new row.
  • The log gains at least one dated row after every run.
  • Every gate writes concrete numbers next to its verdict, never a bare PASS.
  • A hand-off note sits next to the output, listing the checks the second agent re-ran.

Requirements

  • An agent that reads a Markdown instruction file (Claude Code, Codex, Cursor or similar).
  • A bash shell with grep to run log-gate.sh.
  • A recurring workflow with stages you can check.
  • A way to start a second agent session for the hand-off check.

Questions

Why read the whole log instead of searching it?

A search finds only what you ask for. A full read shows the agent rules it did not know to ask about. When the log gets too long, move the rules that gates now enforce into a short list at the top.

What goes in the log when a run needed no corrections?

One row: symptom none, rule no change. The close stays honest, and the log records that the run happened.

Why a second agent if the first one already checked?

An agent's report is a claim, not a check. A second agent that did not build the output re-runs the cheap decisive checks on the raw output. That catches the cases where the first agent checked something other than what it thought.

How do I know a gate works at all?

Make a deliberately broken copy of a real output and run the gate on it. The gate must fail. A gate that passes a broken copy protects nothing.

Can several workflows share one log?

Keep them separate. Rules from a report rarely apply to a deploy, and an agent skims a mixed log instead of reading it.

Where it fits

This template adds memory to every other workflow in this hub. With the orchestrator template, each workstream keeps its own log. With the adversarial review template, review findings land in the log as rows with source QA. With the clarify-before-acting template, answers to questions that keep coming back become rules, so the agent does not ask twice.

All skills

Sources

Questions about setting these up go in the Discord.Join free