Research agent with a source on every claim

The agent answers a research question only once every claim in its ledger points to the page it came from.

Research, beginner. Published

What it does

The agent starts with three lines: the question, the decision that depends on it, and the scope. It then plans its sources and marks up front which it can open itself (opened) and which you have to supply (supplied) because they need a login or are video. Every fact becomes its own ledger row with the page address, the read date, the access type and the scope it holds for. A new row may cite only a page opened or supplied in this session, never an address from memory. A row from an earlier session counts only inside the refresh window; an older one is re-opened or marked needs re-read. The agent's own conclusions are separate rows marked reading. When a page will not open, it sends you a short request: what to open, what to look for, what to send back. Anything missing goes on a gap list. The answer comes last, and every sentence in it names the ledger rows it rests on.

When to use it

When to use it

  • You are choosing a vendor, tool or service and want a decision based on checked pages.
  • You are comparing options and need to know where every number came from.
  • Some sources need a login and you want a clear split of work with the agent.
  • Research runs over several sessions and the findings should stay in a file.

When not to use it

  • A question about knowledge that does not change over time.
  • A brainstorm where ideas matter and evidence does not.
  • Your agent cannot open web pages. Then you have to supply every source yourself.

Decision table

SituationWhat the skill does
The page opensWrites a row with the address, read date and opened access
The page needs a login or is a videoSends you a request with a link and what to send back
The data is not publicRecords a gap with the reason, no guessing
A quote in another languageKeeps the original and adds the explanation as a gloss
The agent's own conclusionRecords it as a reading with the row numbers it rests on
A sentence in the answer with no rowRemoves it or moves it to the gap list

Template

The method: framing the question, planning sources, the claim ledger, requests for material, the gap list and an answer built only from the ledger.

SKILL.mdDownloadSKILL.md
---
name: sourced-research-agent
description: Answer a research question from the web with a claim ledger behind the answer, where each claim records the page it rests on, the day it was read, and whether the agent opened the page or the user supplied it, and where every hole is listed as a gap. Use it when someone asks to "research", "compare", "look into" or "find out" something they will act on, such as choosing a vendor, checking a competitor or sizing up an option.
---

# Research with a claim ledger

When an agent answers a research question in prose, three kinds of statement end up looking identical:
things it just read, things it remembers from training, and things it filled in because they sounded
right. This skill keeps them apart. The agent builds a ledger of claims first, each one tied to the page
it came from, and writes the answer only from that ledger. Whatever the ledger cannot support goes on a
gap list the reader can see.

Before starting, read `CONTEXT.md` next to this file: the question, the scope, the options or competitors
to compare, the sources the reader expects, and where the ledger is kept.

If `CONTEXT.md` is missing, or a field still holds `___`, ask the reader for it before searching, in the
language they write in. For a comparison, ask for the list by name, for example: „Dodaj swoją listę opcji
albo konkurentów do porównania: …" or "Add your list of options or competitors to compare: …". Never
build that list from memory. If the reader wants you to find candidates, record each one you find as a
ledger row with its page, like any other claim.

## Before you start / What you need

- **An AI agent with web search and page fetch**: Claude Code (https://code.claude.com/docs/en/overview)
  has both built in; Codex CLI (https://developers.openai.com/codex/cli) and Cursor
  (https://cursor.com/docs) work too. Without them every source is `supplied`.
- **A file the agent can write**, for the ledger between sessions.
- **Optional: Playwright MCP** (Apache-2.0, https://github.com/microsoft/playwright-mcp) for pages that
  only render in a real browser. Needs Node.js. Claude Code: `claude mcp add playwright npx
  @playwright/mcp@latest`; Codex: `codex mcp add playwright npx "@playwright/mcp@latest"`. Do not sign in
  anywhere with it; a page behind a login stays `supplied`.
- **Optional: the official X API** (https://docs.x.com/x-api/introduction) when public posts are a
  source. The reader creates an app in the Developer Console (https://console.x.com), buys pay-per-use
  credits, sets a spending limit, and generates an app-only token kept in an environment variable in their
  shell profile or a `.env` file listed in `.gitignore`, never in the chat or in `CONTEXT.md`.

If a tool is missing, say so once and mark those sources `supplied`.

## 1. Frame the question

Write three lines at the top of the ledger before any searching:

- **Question:** one sentence.
- **Decision:** what the reader will do differently depending on the answer.
- **Scope:** the place, time period and audience the answer must hold for.

If the question is too broad to answer in one session, split it and say which part this session covers.

## 2. Plan the sources

List the sources you intend to use, starting with the pages that can settle the deal-breakers in
`CONTEXT.md` for each option. For each source, note what it can tell you and how you expect to reach it:

| Access | Meaning |
|---|---|
| `opened` | you can load the page yourself and read it |
| `supplied` | you cannot load it (login, paywall, heavy scripts, video, rate limits), so the reader will open it and paste or upload what you need |

Be honest about your tools. You cannot watch a video, sign in, use the reader's own search engine or see
content that appears only after scripts run in a browser. A source that needs any of these is a
`supplied` source from the start.

## 3. Build the ledger

Every fact you might rely on becomes one ledger row:

```
C<n> | <the claim, one sentence> | <page address> | read <YYYY-MM-DD> | <opened | supplied> | scope: <place / period / audience>
```

Rules for a row:

- **Only pages actually read.** A new row needs an address you loaded in this session, or the source of
  material the reader gave you in it. An address you reconstructed, guessed or remember does not qualify,
  however plausible it looks. Rows carried over from an earlier session follow section 7.
- **Deal-breakers first.** Write the deal-breaker rows for each option before any other rows. Once a
  row shows that an option hits a deal-breaker, that option is ruled out, citing that row, and gets no
  further research. If `CONTEXT.md` says `none`, skip this step.
- **Nothing from memory.** Background knowledge can help you decide where to look. It never becomes a row.
- **Scope is part of the claim.** A price seen on one region's page, a review count on one site, an
  opinion from one forum: record exactly where it holds, so the answer does not stretch it further.
- **Read date, not publish date.** Record when you read the page. If the page states its own date, add it
  inside the claim ("pricing page dated <date>").
- **Quotes are copied, not reworded.** Put the original words in quotation marks, in the language they
  were written in. If you need to explain them in another language, add a `gloss:` after the row and never
  quote the gloss.
- **Facts and readings are separate rows.** A row is either what the page says, or your reading of it,
  marked `reading:`. A reading row cites the fact rows it rests on (`reading: based on C4, C7`).
- **Estimates stay estimates.** Numbers from tools that estimate things they cannot observe (traffic,
  revenue, audience size) are recorded with `third-party estimate` in the claim.
- **Counts must be countable.** If you say "several users report X", list those rows. The number you
  state is the number of rows you list.

## 4. When a source will not open

Do not work around a blocked source by filling it in. Write a short request for the reader instead:

```
REQUEST R<n>
Open: <address, marked "link for you, not evidence">
Look for: <exactly what to check>
Send back: <what to paste or upload>
```

Keep working on everything else while the request is open. When the material arrives, add rows with
access `supplied`. A request nobody answered goes on the gap list.

## 5. Record gaps as you go

A gap is anything the answer needs that the ledger does not hold:

```
GAP G<n> | <what is missing> | why: <blocked, not public, contradictory sources, out of scope> | would close it: <source or action>
```

Some things cannot be seen from outside at all, such as another organisation's internal costs, margins or
private metrics. List them as gaps with `why: not public`, rather than inferring them.

## 6. Write the answer

Only now write the answer, in this order:

1. **Answer:** two to five sentences that respond to the question and the decision. Every factual
   statement ends with the row numbers that support it, like `(C3, C8)`.
2. **Options or ranking,** if the question compares things. Each line names its supporting rows. A line
   that rests only on `reading:` rows says so.
3. **What would change the answer:** the one or two facts that, if different, would flip the
   recommendation.
4. **Gaps:** the full gap list and any open requests. If there are none, write "No gaps recorded" so the
   reader knows the list was checked.

Try to disprove the option you favour before recommending it. If you looked for a reason it fails and
found none, say what you looked for.

## 7. Keep it across sessions

Keep the ledger in the file set in `CONTEXT.md`. Start each session by reading it and noting which rows are
older than the refresh window. A carried-over row may support today's answer only if its read date is
inside the refresh window. A row older than that is either re-opened this session (add a new row and note
`replaces C<n>`) or marked `needs re-read`; a `needs re-read` row cannot support the answer and goes on the
gap list. Never edit an old row in place.

## 8. Checks before sending

- Does every factual sentence in the answer point to at least one ledger row?
- Was every cited row read in this session, or carried over from an earlier one and still inside the
  refresh window?
- Is anything in the answer drawn from memory rather than the ledger?
- Are quotes in their original words and language?
- Does every stated count match the rows listed?
- Is the gap list present?

Fix anything that fails before replying.

## Failure modes

| What you see | What went wrong | What to do |
|---|---|---|
| A figure nobody can trace | the row cites a guessed or remembered address | cite only pages actually loaded or supplied; carried rows only inside the refresh window |
| A confident answer on thin evidence | readings written up as facts | separate `reading:` rows, list gaps |
| One site's data presented as the whole picture | scope left off the row | record where each claim holds |
| The agent "watched" or "searched" something it cannot | tool limits not stated | mark the source `supplied` and send a request |
| Old pages treated as current | read dates ignored between sessions | cite carried rows only inside the refresh window; re-read or mark `needs re-read` |

Other files

CONTEXT.template.mdYour settings: question and decision, scope, expected sources with their access, refresh window, deal-breakers.
CONTEXT.template.mdDownloadCONTEXT.template.md
# Research ledger: your context

Copy this file to CONTEXT.md next to SKILL.md (Claude Code), or next to your AGENTS.md or rules file
(Codex, Cursor), and fill in every blank. Write `no data` for anything you do
not know yet. Never put passwords, access keys or card numbers here.

## Question and decision

(What you want to know, and what you will do with the answer.)
Question: ___
Decision: ___
Example (not your data): Which of three hosting providers fits a small web app with users in one region? Decision: which one to sign up with this month.

## Scope

(The place, time period and audience the answer must hold for.)
Scope: ___
Example (not your data): users in one region, prices as of this month, a team of two developers.

## Options or competitors to compare

(Add your own list. The agent asks for it if it is empty and never makes one up. Write `find candidates`
if you want the agent to search for them, or `none` if the question is not a comparison.)
Options: ___
Example (not your data): three hosting providers we already shortlisted, named here.

## Expected sources

(Where you expect the evidence to be, and whether the agent can open it or you will supply it.)
| Source | What it should tell you | Access |
|---|---|---|
| ___ | ___ | opened / supplied |
Example (not your data): each provider's pricing and documentation pages, opened; a review site that needs a login, supplied; two public developer forums, opened.

## Refresh window

(After how long a ledger row should be re-read before the answer relies on it.)
Refresh window: ______
Example (not your data): 30 days for prices and status pages, 180 days for documentation.

## Deal-breakers

(Facts that rule an option out at once. The agent checks these first. `none` is a valid answer.)
Deal-breakers: ______
Example (not your data): no server location in our region; no way to export our data; price above our monthly budget.

## Ledger location

(The file where claims, requests and gaps are kept between sessions.)
Path: ___
Example (not your data): ./research/ledger.md
ledger.template.mdThe ledger file template: sources, claim rows, requests to you, gaps and session notes.
ledger.template.mdDownloadledger.template.md
# Research ledger

Question: (one sentence)
Decision: (what changes depending on the answer)
Scope: (place / period / audience)

## Sources planned

| Source | What it should tell us | Access |
|---|---|---|
| (name) | (what) | opened / supplied |

## Claims

C1 | (claim, one sentence) | (page address) | read YYYY-MM-DD | opened | scope: (where it holds)
C2 | "(quote in original words)" | (page address) | read YYYY-MM-DD | supplied | scope: (...)
gloss: (explanation in another language, never quoted)
C3 | reading: (what C1 and C2 suggest) | based on C1, C2

## Requests

REQUEST R1
Open: (address, link for you, not evidence)
Look for: (what to check)
Send back: (what to paste or upload)
Status: open / answered on YYYY-MM-DD / unanswered

## Gaps

GAP G1 | (what is missing) | why: (blocked / not public / contradictory / out of scope) | would close it: (source or action)

## Session notes

YYYY-MM-DD: (rows added, rows re-read, rows marked needs re-read, requests sent)

What you need

Install

  1. Claude Code: copy the folder to .claude/skills/sourced-research-agent/ in your project, or to ~/.claude/skills/sourced-research-agent/ for every project. Claude loads the skill when its description matches the task, or when you type /sourced-research-agent.
  2. Codex: paste the contents of SKILL.md into AGENTS.md at the repository root, or into the global ~/.codex/AGENTS.md.
  3. Cursor: add SKILL.md as a project rule, or paste it into AGENTS.md.
  4. A chat with no files: paste SKILL.md at the start of the conversation or into the project instructions.
  5. Copy CONTEXT.template.md to CONTEXT.md next to SKILL.md (Claude Code), or next to your AGENTS.md or rules file (Codex, Cursor), and fill in every field.
  6. Create the ledger file from ledger.template.md at the path set in CONTEXT.md.

It's working if

  • Before searching, the ledger holds the question, the decision and the scope.
  • Every ledger row has a page address, a read date and opened or supplied.
  • When a page will not open, you get a request with a link, what to look for and what to send back.
  • Every factual sentence in the answer ends with row numbers in brackets.
  • The answer ends with a gap list, or a sentence saying no gaps were recorded.

Requirements

  • An agent with a search or page-opening tool.
  • A file where the agent can keep the ledger between sessions.
  • Willingness to open a few blocked pages and send back the result.

Questions

Why can't the agent cite an address from memory?

An address remembered or rebuilt from a pattern looks credible, but nobody checked what is actually there today. A new row accepts only pages opened or supplied in this session, and rows from earlier sessions count only inside the refresh window. So every row can be checked in one click.

Why separate rows for conclusions?

When a fact and an interpretation share a sentence, the interpretation reads like a fact. A reading row shows which facts a conclusion rests on, so you can see at once how strong it is.

What if I don't answer the agent's request?

The agent keeps working on the other sources, and the unanswered request goes on the gap list. The answer states plainly what it is missing.

Doesn't the ledger slow things down?

A little, at first. In return every number in the answer traces to a page and a date, and in the next session the agent knows which rows need re-reading.

Where it fits

This skill is the evidence layer for any research. The voice-of-customer mining template applies a similar discipline to one kind of material: customer words. The parallel fan-out template lets you run several sources at once, each with its own agent, and the gap list works as input for the open-decisions escalation template.

All skills

Sources

Questions about setting these up go in the Discord.Join free