Log poprawek: agent AI, który uczy się na błędach

Twój powtarzalny proces zapamiętuje każdą poprawkę i nie zamknie uruchomienia, dopóki jej nie zapisze.

Operacje, średni poziom. Opublikowano

Co robi

Agent przed każdym uruchomieniem procesu czyta cały log poprawek, od pierwszego do ostatniego wiersza. Każdy etap kończy bramka, która zapisuje zmierzone liczby, a stan "nie udało się zmierzyć" zatrzymuje pracę tak samo jak błąd. Zanim wynik trafi do Ciebie, drugi agent, który go nie budował, sam powtarza najtańsze rozstrzygające testy. Na końcu agent dopisuje do logu każdą poprawkę: objaw, przyczynę i regułę na przyszłość. Skrypt log-gate.sh odmawia zamknięcia uruchomienia, jeśli log nie urósł. Gdy ta sama poprawka pojawia się drugi raz, agent przenosi ją do bramki albo do instrukcji procesu, zamiast dopisywać kolejny wiersz.

Kiedy używać

Kiedy używać

  • Ten sam proces agent wykonuje regularnie: raport, publikację, wdrożenie, montaż, eksport danych.
  • Poprawiasz agenta w tych samych miejscach przy kolejnych uruchomieniach.
  • Chcesz, żeby wynik był sprawdzony, zanim go zobaczysz, a nie po.

Kiedy nie używać

  • Zadanie jednorazowe, które się nie powtórzy.
  • Proces, w którym nikt ani nic nie potrafi wskazać, co poszło źle. Najpierw dodaj bramkę.
  • Preferencje, które zmieniasz przy każdym uruchomieniu.

Tabela decyzyjna

SytuacjaCo robi skill
Poprawiasz wynik po raz pierwszyAgent dopisuje wiersz: objaw, przyczyna, reguła dla całej klasy problemów
Ta sama poprawka wraca drugi razAgent przenosi regułę do bramki albo do instrukcji i zapisuje to w logu
Bramka nie może czegoś zmierzyćEtap staje, a brak pomiaru trafia do logu jako osobny wiersz
Agent twierdzi, że wszystko przeszłoDrugi agent sam powtarza najtańsze rozstrzygające testy przed przekazaniem
Uruchomienie trwa dłużej niż poprzednieDziennik uruchomień pokazuje wolniejszy etap, a agent nazywa przyczynę

Szablon

Metoda: czytanie logu przed każdym uruchomieniem, bramki z liczbami, podwójna kontrola przy przekazaniu, dopisywanie poprawek przed zamknięciem.

SKILL.mdPobierzSKILL.md
---
name: correction-log
description: Makes any recurring agent workflow (a report, an edit, a deploy, a content pipeline) learn from its corrections through an append-only log the agent reads before every run and must grow before the run can close. Use it when the same workflow runs again and again and a human keeps correcting the same kinds of mistakes.
---

# Correction log: a workflow that learns from its mistakes

A recurring workflow repeats its mistakes unless something carries the lesson from one run to
the next. This skill gives the workflow that memory: one append-only log per workflow, read end to
end before every run, written to before every run closes, and backed by gates that measure instead
of assert.

Read `CONTEXT.md` next to this file first. It names the workflow, where its log lives, who counts
as a correction source, the gates, and who does the independent re-check at hand-off. If
`CONTEXT.md` is missing or a field you need is blank, ask the user once, in the language they write
in, then write the answer into `CONTEXT.md` so nobody asks again. For example:

- English: "Add your workflow's stages and the gate that ends each one."
- Polish: „Dodaj etapy swojego procesu i bramkę, która kończy każdy z nich."

The template ships with no workflow, thresholds or owner. Never invent them.

## Before you start / What you need

- **An AI coding agent** that reads a markdown instruction file and can run shell commands: Claude
  Code, Codex CLI or Cursor.
- **A bash shell with grep** to run `log-gate.sh`. macOS and Linux have both. On Windows, use WSL.
- **A second agent session** for the hand-off check: a new session of the same agent is enough, as
  long as it did not build the output.
- **A filled-in `CONTEXT.md`** and a `CORRECTION-LOG.md` copied from the template.

## The four files

- `SKILL.md` (this file): the method. It changes rarely.
- `CONTEXT.md`: the operator's details. It changes when the workflow changes.
- `CORRECTION-LOG.md`: the workflow's memory. It only ever grows. Start from
  `CORRECTION-LOG.template.md` and delete its two example rows.
- `log-gate.sh`: a tiny script that refuses to close a run unless the log grew.

One log per workflow. A weekly report and a deploy pipeline get separate logs, because a rule
learned on one rarely applies to the other and a mixed log gets skimmed instead of read.

## Before the run: read the whole log

1. Open the workflow's `CORRECTION-LOG.md` and read every row, top to bottom. Do not search it
   for keywords and do not read only the latest rows. The log holds mistakes already paid for;
   the point of the read is to avoid paying for one of them twice.
2. Record the baseline so the close check can prove the log grew:

   ```bash
   ./log-gate.sh begin path/to/CORRECTION-LOG.md
   ```

3. If `CONTEXT.md` lists a run ledger, read its last three rows. They set the time you expect
   each stage to take on this run.
4. Before you start, state in one line which log rules apply to this run. A rule you cannot
   connect to this run still gets checked in the hand-off log re-read below; the line only proves the
   read happened.

## During the run: gates measure, they do not assert

Each stage of the workflow ends in a gate: a check that measures the thing that went wrong
before. A stage is finished when its gate passes with numbers written down, not when the output
looks done.

Every gate follows three rules:

1. **Write numbers or nothing.** A gate records what it measured (row counts, totals, durations,
   diffs, error counts) next to its verdict. A bare "PASS" is not evidence. If the gate cannot
   produce the number, it does not produce a verdict either.
2. **"Could not measure" is its own exit state.** Use three outcomes: pass, fail, and could not
   measure. Could not measure halts the stage just like a fail; it is never rounded up to a pass.
   A missing input, a timeout, or a tool that returned nothing all land here.
3. **Watch it fail before you trust it.** Before a gate goes into service, and after any change
   to it, run it against a deliberately broken copy of a real output: a total that no longer
   adds up, a missing section, a truncated file, whatever the gate claims to catch. The gate must
   fail on that copy. A gate that passes a broken copy is decoration. Write the broken-copy test
   down in `CONTEXT.md` so it can be repeated.

If a gate fails twice on the same class of problem in one run, stop. Report what the gate
measured, which numbers sit out of tolerance, the options, and the one you recommend. Do not
lower a threshold to reach green; a threshold moves only when the owner named in `CONTEXT.md`
says so, and that change is itself a log row.

## Hand-off: separation of duty

Nothing reaches the human until the hand-off check passes and its results are written to a
hand-off note next to the output.

**The builder** (the agent that produced the output) runs every gate in `CONTEXT.md` plus:

- **Log re-read.** Walk the correction log against the finished output, row by row. Each
  past rule is checked against this output, not assumed to still hold.

**The checker** (a second agent that did not build the output, named in `CONTEXT.md`) re-runs
the cheap decisive checks itself before the human sees anything. Cheap means seconds, not
minutes. Decisive means a failure there would make the output wrong, not merely less polished.
`CONTEXT.md` lists which checks those are.

The builder's report does not count as the check. "The agent said it passed" is a claim, not a
check. If the checker cannot run a check, the output does not ship; the hand-off note says which
check could not run and why.

## After the run: write corrections back before closing

The run does not close until the log grew. Append one row per lesson, in this shape:

```
| Date | Source | Correction |
| YYYY-MM-DD | user / QA / post-mortem | Symptom: ... Cause: ... Rule: ... |
```

- **Source** says where the correction came from: `user` (the human corrected the output), `QA`
  (a gate or the checker caught it), or `post-mortem` (found after delivery).
- **Symptom** is what was visibly wrong. **Cause** is why it happened. **Rule** is the
  instruction that now applies to every future run, written as an imperative a new agent could
  follow with no other context.

What counts as a row:

- every mistake the human corrected, however small;
- every workaround you had to invent;
- every threshold the owner moved;
- every gate that could not measure;
- every gate that let a broken copy through.

Then close the run:

```bash
./log-gate.sh close path/to/CORRECTION-LOG.md
```

It exits non-zero unless the log has more rows than at `begin`. A run with genuinely nothing to
learn still adds a row: `Symptom: none. Rule: no change.` That row keeps the close honest and
records that the run happened.

## Fix the class, not the instance

When you write the rule column, name the class of problem, not the single case. "Last month's total
was wrong" fixes one report. "Totals are recomputed from line items and compared with the
source's own total before the report is written" fixes every future report. If you can only
write the instance, the cause column is not finished yet.

## A correction given twice is a pipeline failure

Before you append a row, search the log for the same class. If an earlier row already covers it,
the log did not work for that lesson, and another row will not work either. Instead:

1. Turn the rule into a gate: a check that measures it on every run, with a broken-copy test.
2. Or move it into the body of the skill or workflow instructions, where it is followed rather
   than remembered.
3. Then append a row that says so: `Rule: repeat of <earlier date>; now enforced by <gate or
   section>.`

A log where the same class appears three times is a sign the escalation step is being skipped.

## Optional: the run ledger

If `CONTEXT.md` enables it, append one row per run to a separate `RUN-LEDGER.md`:

```
| Date | Run | Stage timings | Log rules applied | Rows added | Note |
```

Record wall time per stage from the start of the stage to its gate passing. Write `no data` for
anything you did not measure; never estimate. Read the last three rows before each run. A stage
that runs clearly slower than the previous comparable run is a finding: name the cause, and if it
repeats, it becomes a correction-log row. Over time the ledger shows whether the workflow is
getting faster and cleaner, or only growing a longer log.

## Maintaining the log

- **Append only.** Never edit or delete a past row. If a rule turns out wrong, append a row that
  supersedes it and names the earlier date.
- **Keep rows short.** One to three sentences. The log is read end to end every run, so padding
  costs reading time on every run.
- **Graduate rules.** When a rule has become a gate or a section of the instructions, leave the
  row in place and note the promotion in a new row. The history stays; the enforcement moves.
- **Split when a log grows too long to read in one sitting.** Move rules that are now enforced by
  gates into a short "enforced elsewhere" list at the top, and keep the table for live lessons.

## Failure modes to watch for

| What you see | What it means | What to do |
|---|---|---|
| The run closes without a new row | The close check was skipped | Run `log-gate.sh close` as the last step, every run |
| Rows describe one case each | The class was never named | Rewrite the rule column as a class before closing |
| The same class appears twice | The log is not changing behaviour | Promote the rule to a gate or the instructions |
| A gate only ever prints PASS | It records no numbers or was never tested | Add the numbers and run the broken-copy test |
| The checker always agrees with the builder | The checker is reading the builder's report | Make the checker re-run the checks from the raw output |
| Thresholds drift over time | Someone lowered a bar to get green | Only the named owner moves a threshold, and it is logged |

## When not to use this

- A one-off task that will not run again. There is no next run to learn.
- A workflow with no human or gate that can say what went wrong. Add a gate first.
- Personal taste that changes every run. Log only corrections that should still hold next time.

Pozostałe pliki

CORRECTION-LOG.template.mdPusty log poprawek z nagłówkiem i dwoma przykładowymi wierszami do usunięcia.
CORRECTION-LOG.template.mdPobierzCORRECTION-LOG.template.md
# Correction log: [workflow name]

Append-only. Read every row before each run; append before the run closes.
A correction given twice is a pipeline failure: promote it to a gate or the instructions.

Each row: date · source (user / QA / post-mortem) · symptom, cause, and the rule that now applies.
Never edit or delete a past row; supersede it with a new one that names the earlier date.

## Enforced elsewhere

(Rules that have become gates or instruction sections. One line each, with the gate or section name.)

## Corrections

| Date | Source | Correction |
|---|---|---|
| 2000-01-01 | user | Example row, delete it. Symptom: the weekly sales report listed refunded orders as revenue. Cause: the query read gross order totals and ignored the refund table. Rule: revenue is net of refunds; the report gate compares the net total with the payment provider's payout total and fails on any gap. |
| 2000-01-08 | QA | Example row, delete it. Symptom: the report shipped with an empty chart. Cause: the export step timed out and the chart step read a zero-byte file as "no sales". Rule: a zero-byte or missing input is "could not measure" and halts the run; it is never read as zero. |
log-gate.shMały skrypt bash: begin zapisuje liczbę wierszy, close kończy się błędem, jeśli log nie urósł.
log-gate.shPobierzlog-gate.sh
#!/usr/bin/env bash
# log-gate.sh: refuse to close a run unless its correction log grew.
#   log-gate.sh begin <log>   record the current row count in <log>.baseline
#   log-gate.sh close <log>   exit 0 only if the log now has more rows than the baseline
# A row is a table line that starts with a date: "| YYYY-MM-DD |".
# Exit: 0 ok · 1 log did not grow · 2 could not measure (missing log or baseline) · 3 usage.
set -u

count_rows() { grep -cE '^\|[[:space:]]*[0-9]{4}-[0-9]{2}-[0-9]{2}[[:space:]]*\|' "$1"; }

[ $# -eq 2 ] || { echo "usage: log-gate.sh begin|close <log>" >&2; exit 3; }
cmd="$1"; log="$2"; base="$log.baseline"
[ -f "$log" ] || { echo "could not measure: log not found: $log" >&2; exit 2; }
now="$(count_rows "$log")"

case "$cmd" in
  begin)
    echo "$now" > "$base" || { echo "could not measure: cannot write $base" >&2; exit 2; }
    echo "begin: $now rows recorded in $base"
    ;;
  close)
    [ -f "$base" ] || { echo "could not measure: no baseline, run 'begin' first" >&2; exit 2; }
    was="$(cat "$base")"
    case "$was" in ''|*[!0-9]*) echo "could not measure: bad baseline in $base" >&2; exit 2;; esac
    if [ "$now" -gt "$was" ]; then
      echo "close: log grew from $was to $now rows"
      rm -f "$base"
    else
      echo "close refused: log has $now rows, baseline $was. Append this run's corrections first." >&2
      exit 1
    fi
    ;;
  *) echo "usage: log-gate.sh begin|close <log>" >&2; exit 3 ;;
esac
CONTEXT.template.mdSzablon kontekstu: Twój proces, etapy, bramki, progi, osoba kontrolująca i dziennik uruchomień.
CONTEXT.template.mdPobierzCONTEXT.template.md
# Correction log: context

Copy this file to CONTEXT.md next to SKILL.md and fill in every blank (______). Example lines only show the shape of an answer; delete them once you have written your own. Never put passwords, access keys or card numbers here.

## The workflow

(What recurring workflow this log serves, how often it runs, and what it delivers.)
Your answer: ______
Example (replace it, this is not your data): Weekly sales report, at the start of each week, a one-page summary sent to the shop owner.

## Stages

(The stages in order, each with the gate that ends it. One line per stage.)
Your answer: ______
Example (replace it, this is not your data): 1. Pull orders and refunds. Gate: row counts match the source export.
Example (replace it, this is not your data): 2. Compute totals. Gate: net total matches the payment provider's payout within the tolerance below.
Example (replace it, this is not your data): 3. Write the report. Gate: every section present, no placeholder text.

## Where the files live

(Path to this workflow's CORRECTION-LOG.md, to log-gate.sh, and to the run ledger if you use one.)
Your answer: ______
Example (replace it, this is not your data): reports/weekly-sales/CORRECTION-LOG.md

## Correction sources

(Who or what may add a row: the people whose corrections count, the gates that count as QA, and when a post-mortem happens.)
Your answer: ______
Example (replace it, this is not your data): user = the shop owner; QA = any gate failure or checker finding; post-mortem = any error found after the report was sent.

## Thresholds and their owner

(Each gate threshold, its value, how you derived it, and the one person allowed to change it.)
Your answer: ______
Example (replace it, this is not your data): payout match tolerance: the smallest gap your payment provider's rounding produces in a normal week. Owner: the shop owner.

## Broken-copy tests

(For each gate, the deliberately broken copy it must fail on, and when you last watched it fail.)
Your answer: ______
Example (replace it, this is not your data): totals gate: copy last week's report, delete one refund row, confirm the gate fails.

## Hand-off

(Who the checker is, which cheap decisive checks the checker re-runs, and where the hand-off note is written.)
Your answer: ______
Example (replace it, this is not your data): checker = a second agent session that did not build the report; it re-runs the totals gate and the section check from the raw export.

## Run ledger

(On or off. If on, the stages you time.)
Your answer: ______
Example (replace it, this is not your data): on; time each of the three stages above.

## Notes for the agent

(Anything else a new agent needs before its first run: tone of the output, who receives it, known quirks of the data.)
Your answer: ______

Czego potrzebujesz

  • An AI coding agent: Claude Code, Codex CLI or Cursor (otwiera się w nowej karcie)

    Wymagane · Aplikacja

    Prowadzi proces, czyta log i uruchamia log-gate.sh. Potrzebny jest agent, który czyta plik instrukcji w markdown i może uruchamiać polecenia w powłoce.

    1. Zainstaluj jednego agenta według jego oficjalnej dokumentacji: Claude Code (code.claude.com/docs), Codex CLI (developers.openai.com/codex) albo Cursor (cursor.com/docs).
    2. Umieść folder szablonu tam, skąd Twój agent wczytuje skille lub instrukcje, zgodnie z krokami instalacji.
    3. Do kontroli przy przekazaniu otwórz drugą, nową sesję tego samego agenta. Nie może to być sesja, która zbudowała wynik.
  • Bash shell with grep (otwiera się w nowej karcie)

    Wymagane · Środowisko · Licencja GPL-3.0-or-later

    log-gate.sh to skrypt bash, który liczy wiersze logu za pomocą grep.

    1. macOS i Linux mają bash i grep od razu. Sprawdź to poleceniem: bash --version.
    2. Windows: zainstaluj WSL według instrukcji Microsoft (learn.microsoft.com/windows/wsl/install) i uruchamiaj skrypt w WSL.
    3. Nadaj skryptowi prawo wykonania: chmod +x log-gate.sh.

Instalacja

  1. Claude Code: skopiuj folder do .claude/skills/correction-log/ w projekcie albo do ~/.claude/skills/correction-log/ dla wszystkich projektów. Claude wczyta go, gdy opis pasuje do zadania, albo gdy wpiszesz /correction-log.
  2. Codex: wklej treść SKILL.md do pliku AGENTS.md w katalogu głównym repozytorium albo w ~/.codex/AGENTS.md.
  3. Cursor: dodaj treść SKILL.md jako regułę projektu albo do AGENTS.md.
  4. Inny agent: wklej SKILL.md do pliku instrukcji projektu lub promptu systemowego.
  5. Skopiuj CONTEXT.template.md jako CONTEXT.md obok SKILL.md i wypełnij każde pole.
  6. Skopiuj CORRECTION-LOG.template.md obok swojego procesu jako CORRECTION-LOG.md, usuń dwa przykładowe wiersze i nadaj log-gate.sh prawo wykonania (chmod +x).

Działa, jeśli

  • Pierwsza odpowiedź agenta w uruchomieniu wymienia reguły z logu, które dotyczą tego uruchomienia.
  • Obok logu pojawia się plik .baseline po starcie uruchomienia i znika po poprawnym zamknięciu.
  • log-gate.sh close kończy się kodem 1, gdy spróbujesz zamknąć uruchomienie bez nowego wiersza.
  • Log zyskuje co najmniej jeden datowany wiersz po każdym uruchomieniu.
  • Każda bramka zapisuje obok werdyktu konkretne liczby, a nie samo PASS.
  • Obok wyniku leży notatka przekazania z testami powtórzonymi przez drugiego agenta.

Wymagania

  • Agent, który czyta plik instrukcji w Markdown (Claude Code, Codex, Cursor lub podobny).
  • Powłoka bash z grep, żeby uruchomić log-gate.sh.
  • Powtarzalny proces z etapami, które da się sprawdzić.
  • Możliwość uruchomienia drugiej sesji agenta do kontroli przy przekazaniu.

Pytania

Po co czytać cały log, skoro można go przeszukać?

Wyszukiwanie znajduje tylko to, o co pytasz. Pełne czytanie pokazuje agentowi reguły, o które nie wiedział, że powinien zapytać. Gdy log robi się za długi, przenieś reguły wymuszone przez bramki na krótką listę na górze.

Co wpisać, jeśli uruchomienie przeszło bez żadnej poprawki?

Jeden wiersz: objaw brak, reguła bez zmian. Dzięki temu zamknięcie pozostaje uczciwe, a log pokazuje, że uruchomienie się odbyło.

Dlaczego drugi agent, skoro pierwszy już sprawdził?

Raport agenta to deklaracja, a nie test. Drugi agent, który nie budował wyniku, powtarza najtańsze rozstrzygające testy na surowym wyniku. Tak wyłapiesz przypadki, w których pierwszy agent sprawdził coś innego, niż myślał.

Jak sprawdzić, czy bramka w ogóle działa?

Przygotuj celowo zepsutą kopię prawdziwego wyniku i uruchom na niej bramkę. Musi się zatrzymać. Bramka, która przepuszcza zepsutą kopię, niczego nie chroni.

Czy mogę mieć jeden log dla kilku procesów?

Lepiej nie. Reguły z raportu rzadko pasują do wdrożenia, a wspólny log agent zaczyna przeglądać pobieżnie zamiast czytać.

Gdzie to pasuje

Ten szablon dokłada pamięć do każdego innego procesu w tym katalogu. Z szablonem orkiestratora dostajesz osobny log dla każdego strumienia pracy. Z szablonem przeglądu krytycznego wnioski z przeglądu trafiają do logu jako wiersze ze źródłem QA. Z szablonem doprecyzowania przed działaniem odpowiedzi na pytania, które wracają, zamieniasz w reguły, żeby agent nie pytał drugi raz.

Wszystkie skille

Źródła

Pytania o wdrożenie zadajesz na Discordzie.Dołącz za darmo