Agent AI do researchu ze źródłem przy każdym twierdzeniu

Agent odpowiada na pytanie z researchu dopiero wtedy, gdy każde twierdzenie ma w rejestrze stronę, z której pochodzi.

Research, dla początkujących. Opublikowano

Co robi

Agent zaczyna od trzech linijek: pytanie, decyzja, która od niego zależy, i zakres. Potem planuje źródła i od razu oznacza, które otworzy sam (opened), a które musisz mu dostarczyć (supplied), bo wymagają logowania albo są wideo. Każdy fakt trafia do rejestru jako osobny wiersz z adresem strony, datą odczytu, sposobem dostępu i zakresem, w którym obowiązuje. Nowy wiersz może wskazywać tylko stronę otwartą albo dostarczoną w tej sesji, nigdy adres z pamięci. Wiersz z wcześniejszej sesji liczy się tylko w oknie odświeżania; starszy agent otwiera ponownie albo oznacza jako needs re-read. Własne wnioski agent zapisuje jako osobne wiersze z oznaczeniem reading. Gdy strona się nie otwiera, wysyła ci krótką prośbę: co otworzyć, czego szukać, co odesłać. Braki zapisuje na liście luk. Odpowiedź pisze na końcu, a przy każdym zdaniu podaje numery wierszy, na których się opiera.

Kiedy używać

Kiedy używać

  • Wybierasz dostawcę, narzędzie albo usługę i chcesz decyzji opartej na sprawdzonych stronach.
  • Porównujesz opcje i potrzebujesz wiedzieć, skąd pochodzi każda liczba.
  • Część źródeł wymaga logowania i potrzebujesz jasnego podziału pracy z agentem.
  • Research trwa kilka sesji, a ustalenia mają zostać w pliku.

Kiedy nie używać

  • Pytanie o wiedzę, która nie zmienia się w czasie.
  • Burza mózgów, w której liczą się pomysły, a nie dowody.
  • Twój agent nie potrafi otwierać stron. Wtedy każde źródło musisz dostarczyć sam.

Tabela decyzyjna

SytuacjaCo robi skill
Strona się otwieraZapisuje wiersz z adresem, datą odczytu i dostępem opened
Strona wymaga logowania albo to wideoWysyła ci prośbę z linkiem i opisem, co odesłać
Dane nie są publiczneZapisuje lukę z powodem, bez zgadywania
Cytat w innym językuZostawia oryginał, objaśnienie dopisuje jako gloss
Własny wniosek agentaZapisuje go jako reading z numerami wierszy, na których stoi
Zdanie w odpowiedzi bez wierszaUsuwa je albo przenosi do listy luk

Szablon

Metoda: ramowanie pytania, plan źródeł, rejestr twierdzeń, prośby o materiał, lista luk i odpowiedź zbudowana tylko z rejestru.

SKILL.mdPobierzSKILL.md
---
name: sourced-research-agent
description: Answer a research question from the web with a claim ledger behind the answer, where each claim records the page it rests on, the day it was read, and whether the agent opened the page or the user supplied it, and where every hole is listed as a gap. Use it when someone asks to "research", "compare", "look into" or "find out" something they will act on, such as choosing a vendor, checking a competitor or sizing up an option.
---

# Research with a claim ledger

When an agent answers a research question in prose, three kinds of statement end up looking identical:
things it just read, things it remembers from training, and things it filled in because they sounded
right. This skill keeps them apart. The agent builds a ledger of claims first, each one tied to the page
it came from, and writes the answer only from that ledger. Whatever the ledger cannot support goes on a
gap list the reader can see.

Before starting, read `CONTEXT.md` next to this file: the question, the scope, the options or competitors
to compare, the sources the reader expects, and where the ledger is kept.

If `CONTEXT.md` is missing, or a field still holds `___`, ask the reader for it before searching, in the
language they write in. For a comparison, ask for the list by name, for example: „Dodaj swoją listę opcji
albo konkurentów do porównania: …" or "Add your list of options or competitors to compare: …". Never
build that list from memory. If the reader wants you to find candidates, record each one you find as a
ledger row with its page, like any other claim.

## Before you start / What you need

- **An AI agent with web search and page fetch**: Claude Code (https://code.claude.com/docs/en/overview)
  has both built in; Codex CLI (https://developers.openai.com/codex/cli) and Cursor
  (https://cursor.com/docs) work too. Without them every source is `supplied`.
- **A file the agent can write**, for the ledger between sessions.
- **Optional: Playwright MCP** (Apache-2.0, https://github.com/microsoft/playwright-mcp) for pages that
  only render in a real browser. Needs Node.js. Claude Code: `claude mcp add playwright npx
  @playwright/mcp@latest`; Codex: `codex mcp add playwright npx "@playwright/mcp@latest"`. Do not sign in
  anywhere with it; a page behind a login stays `supplied`.
- **Optional: the official X API** (https://docs.x.com/x-api/introduction) when public posts are a
  source. The reader creates an app in the Developer Console (https://console.x.com), buys pay-per-use
  credits, sets a spending limit, and generates an app-only token kept in an environment variable in their
  shell profile or a `.env` file listed in `.gitignore`, never in the chat or in `CONTEXT.md`.

If a tool is missing, say so once and mark those sources `supplied`.

## 1. Frame the question

Write three lines at the top of the ledger before any searching:

- **Question:** one sentence.
- **Decision:** what the reader will do differently depending on the answer.
- **Scope:** the place, time period and audience the answer must hold for.

If the question is too broad to answer in one session, split it and say which part this session covers.

## 2. Plan the sources

List the sources you intend to use, starting with the pages that can settle the deal-breakers in
`CONTEXT.md` for each option. For each source, note what it can tell you and how you expect to reach it:

| Access | Meaning |
|---|---|
| `opened` | you can load the page yourself and read it |
| `supplied` | you cannot load it (login, paywall, heavy scripts, video, rate limits), so the reader will open it and paste or upload what you need |

Be honest about your tools. You cannot watch a video, sign in, use the reader's own search engine or see
content that appears only after scripts run in a browser. A source that needs any of these is a
`supplied` source from the start.

## 3. Build the ledger

Every fact you might rely on becomes one ledger row:

```
C<n> | <the claim, one sentence> | <page address> | read <YYYY-MM-DD> | <opened | supplied> | scope: <place / period / audience>
```

Rules for a row:

- **Only pages actually read.** A new row needs an address you loaded in this session, or the source of
  material the reader gave you in it. An address you reconstructed, guessed or remember does not qualify,
  however plausible it looks. Rows carried over from an earlier session follow section 7.
- **Deal-breakers first.** Write the deal-breaker rows for each option before any other rows. Once a
  row shows that an option hits a deal-breaker, that option is ruled out, citing that row, and gets no
  further research. If `CONTEXT.md` says `none`, skip this step.
- **Nothing from memory.** Background knowledge can help you decide where to look. It never becomes a row.
- **Scope is part of the claim.** A price seen on one region's page, a review count on one site, an
  opinion from one forum: record exactly where it holds, so the answer does not stretch it further.
- **Read date, not publish date.** Record when you read the page. If the page states its own date, add it
  inside the claim ("pricing page dated <date>").
- **Quotes are copied, not reworded.** Put the original words in quotation marks, in the language they
  were written in. If you need to explain them in another language, add a `gloss:` after the row and never
  quote the gloss.
- **Facts and readings are separate rows.** A row is either what the page says, or your reading of it,
  marked `reading:`. A reading row cites the fact rows it rests on (`reading: based on C4, C7`).
- **Estimates stay estimates.** Numbers from tools that estimate things they cannot observe (traffic,
  revenue, audience size) are recorded with `third-party estimate` in the claim.
- **Counts must be countable.** If you say "several users report X", list those rows. The number you
  state is the number of rows you list.

## 4. When a source will not open

Do not work around a blocked source by filling it in. Write a short request for the reader instead:

```
REQUEST R<n>
Open: <address, marked "link for you, not evidence">
Look for: <exactly what to check>
Send back: <what to paste or upload>
```

Keep working on everything else while the request is open. When the material arrives, add rows with
access `supplied`. A request nobody answered goes on the gap list.

## 5. Record gaps as you go

A gap is anything the answer needs that the ledger does not hold:

```
GAP G<n> | <what is missing> | why: <blocked, not public, contradictory sources, out of scope> | would close it: <source or action>
```

Some things cannot be seen from outside at all, such as another organisation's internal costs, margins or
private metrics. List them as gaps with `why: not public`, rather than inferring them.

## 6. Write the answer

Only now write the answer, in this order:

1. **Answer:** two to five sentences that respond to the question and the decision. Every factual
   statement ends with the row numbers that support it, like `(C3, C8)`.
2. **Options or ranking,** if the question compares things. Each line names its supporting rows. A line
   that rests only on `reading:` rows says so.
3. **What would change the answer:** the one or two facts that, if different, would flip the
   recommendation.
4. **Gaps:** the full gap list and any open requests. If there are none, write "No gaps recorded" so the
   reader knows the list was checked.

Try to disprove the option you favour before recommending it. If you looked for a reason it fails and
found none, say what you looked for.

## 7. Keep it across sessions

Keep the ledger in the file set in `CONTEXT.md`. Start each session by reading it and noting which rows are
older than the refresh window. A carried-over row may support today's answer only if its read date is
inside the refresh window. A row older than that is either re-opened this session (add a new row and note
`replaces C<n>`) or marked `needs re-read`; a `needs re-read` row cannot support the answer and goes on the
gap list. Never edit an old row in place.

## 8. Checks before sending

- Does every factual sentence in the answer point to at least one ledger row?
- Was every cited row read in this session, or carried over from an earlier one and still inside the
  refresh window?
- Is anything in the answer drawn from memory rather than the ledger?
- Are quotes in their original words and language?
- Does every stated count match the rows listed?
- Is the gap list present?

Fix anything that fails before replying.

## Failure modes

| What you see | What went wrong | What to do |
|---|---|---|
| A figure nobody can trace | the row cites a guessed or remembered address | cite only pages actually loaded or supplied; carried rows only inside the refresh window |
| A confident answer on thin evidence | readings written up as facts | separate `reading:` rows, list gaps |
| One site's data presented as the whole picture | scope left off the row | record where each claim holds |
| The agent "watched" or "searched" something it cannot | tool limits not stated | mark the source `supplied` and send a request |
| Old pages treated as current | read dates ignored between sessions | cite carried rows only inside the refresh window; re-read or mark `needs re-read` |

Pozostałe pliki

CONTEXT.template.mdTwoje ustawienia: pytanie i decyzja, zakres, spodziewane źródła z dostępem, okno odświeżania, warunki wykluczające.
CONTEXT.template.mdPobierzCONTEXT.template.md
# Research ledger: your context

Copy this file to CONTEXT.md next to SKILL.md (Claude Code), or next to your AGENTS.md or rules file
(Codex, Cursor), and fill in every blank. Write `no data` for anything you do
not know yet. Never put passwords, access keys or card numbers here.

## Question and decision

(What you want to know, and what you will do with the answer.)
Question: ___
Decision: ___
Example (not your data): Which of three hosting providers fits a small web app with users in one region? Decision: which one to sign up with this month.

## Scope

(The place, time period and audience the answer must hold for.)
Scope: ___
Example (not your data): users in one region, prices as of this month, a team of two developers.

## Options or competitors to compare

(Add your own list. The agent asks for it if it is empty and never makes one up. Write `find candidates`
if you want the agent to search for them, or `none` if the question is not a comparison.)
Options: ___
Example (not your data): three hosting providers we already shortlisted, named here.

## Expected sources

(Where you expect the evidence to be, and whether the agent can open it or you will supply it.)
| Source | What it should tell you | Access |
|---|---|---|
| ___ | ___ | opened / supplied |
Example (not your data): each provider's pricing and documentation pages, opened; a review site that needs a login, supplied; two public developer forums, opened.

## Refresh window

(After how long a ledger row should be re-read before the answer relies on it.)
Refresh window: ______
Example (not your data): 30 days for prices and status pages, 180 days for documentation.

## Deal-breakers

(Facts that rule an option out at once. The agent checks these first. `none` is a valid answer.)
Deal-breakers: ______
Example (not your data): no server location in our region; no way to export our data; price above our monthly budget.

## Ledger location

(The file where claims, requests and gaps are kept between sessions.)
Path: ___
Example (not your data): ./research/ledger.md
ledger.template.mdWzór pliku rejestru: źródła, wiersze twierdzeń, prośby do ciebie, luki i notatki z sesji.
ledger.template.mdPobierzledger.template.md
# Research ledger

Question: (one sentence)
Decision: (what changes depending on the answer)
Scope: (place / period / audience)

## Sources planned

| Source | What it should tell us | Access |
|---|---|---|
| (name) | (what) | opened / supplied |

## Claims

C1 | (claim, one sentence) | (page address) | read YYYY-MM-DD | opened | scope: (where it holds)
C2 | "(quote in original words)" | (page address) | read YYYY-MM-DD | supplied | scope: (...)
gloss: (explanation in another language, never quoted)
C3 | reading: (what C1 and C2 suggest) | based on C1, C2

## Requests

REQUEST R1
Open: (address, link for you, not evidence)
Look for: (what to check)
Send back: (what to paste or upload)
Status: open / answered on YYYY-MM-DD / unanswered

## Gaps

GAP G1 | (what is missing) | why: (blocked / not public / contradictory / out of scope) | would close it: (source or action)

## Session notes

YYYY-MM-DD: (rows added, rows re-read, rows marked needs re-read, requests sent)

Czego potrzebujesz

Instalacja

  1. Claude Code: skopiuj folder do .claude/skills/sourced-research-agent/ w projekcie albo do ~/.claude/skills/sourced-research-agent/ dla wszystkich projektów. Claude ładuje skill, gdy opis pasuje do zadania, albo gdy wpiszesz /sourced-research-agent.
  2. Codex: wklej treść SKILL.md do AGENTS.md w katalogu repozytorium albo do globalnego ~/.codex/AGENTS.md.
  3. Cursor: dodaj SKILL.md jako regułę projektu albo wklej go do AGENTS.md.
  4. Czat bez plików: wklej SKILL.md na początku rozmowy albo do instrukcji projektu.
  5. Skopiuj CONTEXT.template.md jako CONTEXT.md obok SKILL.md (Claude Code) albo obok swojego AGENTS.md lub pliku z regułami (Codex, Cursor) i uzupełnij każde pole.
  6. Utwórz plik rejestru z ledger.template.md w miejscu podanym w CONTEXT.md.

Działa, jeśli

  • Zanim agent zacznie szukać, w rejestrze stoją pytanie, decyzja i zakres.
  • Każdy wiersz rejestru ma adres strony, datę odczytu i opened albo supplied.
  • Gdy strona się nie otwiera, dostajesz prośbę z linkiem, opisem i tym, co masz odesłać.
  • Każde zdanie faktu w odpowiedzi kończy się numerami wierszy w nawiasie.
  • Odpowiedź kończy się listą luk albo zdaniem, że luk nie zapisano.

Wymagania

  • Agent z narzędziem do wyszukiwania albo otwierania stron.
  • Plik, w którym agent może trzymać rejestr między sesjami.
  • Twoja gotowość, by otworzyć kilka zablokowanych stron i odesłać wynik.

Pytania

Dlaczego agent nie może podać adresu z pamięci?

Adres zapamiętany albo odtworzony z wzorca wygląda wiarygodnie, ale nikt nie sprawdził, co naprawdę jest pod nim dziś. Nowy wiersz przyjmuje tylko strony otwarte albo dostarczone w tej sesji, a wiersze z wcześniejszych sesji liczą się tylko w oknie odświeżania. Dzięki temu każdy wiersz da się sprawdzić jednym kliknięciem.

Po co osobne wiersze na wnioski?

Gdy fakt i interpretacja stoją w jednym zdaniu, interpretacja wygląda jak fakt. Wiersz reading pokazuje, na których faktach stoi wniosek, więc od razu widzisz, jak mocny jest.

Co jeśli nie odpowiem na prośbę agenta?

Agent pracuje dalej na pozostałych źródłach, a niespełniona prośba trafia na listę luk. Odpowiedź mówi wprost, czego w niej brakuje.

Czy rejestr nie spowalnia pracy?

Trochę, na początku. W zamian każdą liczbę w odpowiedzi da się prześledzić do strony i daty, a przy kolejnej sesji agent wie, które wiersze trzeba odświeżyć.

Gdzie to pasuje

Ten skill jest warstwą dowodową dla każdego researchu. Szablon analizy opinii klientów stosuje podobną dyscyplinę do jednego materiału: słów klientów. Szablon równoległych sekcji pozwala puścić kilka źródeł naraz, każde z własnym agentem, a lista luk nadaje się jako wejście do szablonu eskalacji otwartych decyzji.

Wszystkie skille

Źródła

Pytania o wdrożenie zadajesz na Discordzie.Dołącz za darmo