---
name: voice-of-customer-mining
description: Collect what customers actually say about a problem or product category from reviews, forums, comments and support messages, keep every quote verbatim with its source, cluster the quotes by pain, desire and objection with honest counts, and turn the clusters into angles and message candidates. Use it before writing copy, planning content, positioning an offer, or when someone asks "what are our customers really saying".
---

# Voice-of-customer mining

Copy written from a summary ("customers care about quality") sounds like every competitor's copy. The
words that land are the ones customers already use: specific, sometimes rough, often misspelled. This
workflow collects those words as evidence, counts them honestly, and only then turns them into angles. It
never invents a quote and never passes off a translation or a paraphrase as the customer's voice.

Read `CONTEXT.md` next to this file first. It names the product or problem, the sources you may mine, the
competitors whose customers you may read, the language that counts as primary, and where the cluster bank
lives.

If `CONTEXT.md` is missing, or a field still holds `___`, ask the user for it before collecting anything,
in the language they write in. Ask for the competitor list and the sources by name, for example:
„Dodaj swoją listę konkurentów i źródeł opinii: …" or "Add your competitor list and the places your
customers talk: …". Never make up a competitor or source list yourself. "No competitors" is a valid
answer; write it down and carry on.

## Before you start / What you need

- **An AI agent that can open web pages**: Claude Code (https://code.claude.com/docs/en/overview) has web
  search and page fetch built in; Codex CLI (https://developers.openai.com/codex/cli) and Cursor
  (https://cursor.com/docs) work too. Without page access the user pastes the content and every quote is
  `supplied`.
- **Optional: Playwright MCP** (Apache-2.0, https://github.com/microsoft/playwright-mcp) for review pages
  that only load in a real browser. Needs Node.js. Claude Code: `claude mcp add playwright npx
  @playwright/mcp@latest`; Codex: `codex mcp add playwright npx "@playwright/mcp@latest"`. Do not sign in
  anywhere with it; a page behind a login stays `supplied`.
- **Optional: the official X API** (https://docs.x.com/x-api/introduction) to read public posts. The user
  creates an app in the Developer Console (https://console.x.com), buys pay-per-use credits, sets a
  spending limit, and generates an app-only token. The token lives in an environment variable in the
  user's shell profile or a `.env` file listed in `.gitignore`, never in the chat or in `CONTEXT.md`. Use
  the recent search endpoint (https://docs.x.com/x-api/posts/search/introduction).
- **Optional: review exports** from the user's own store or review app, stripped of personal data first.

If a tool is missing, say so once and mark those sources `supplied`.

## 1. Collect

For each source in `CONTEXT.md`:

1. Open it and read what is actually there. If it will not load or needs a login, do not guess: give the
   reader a ready link and say exactly what to copy back, then work from what they paste.
2. Copy each relevant statement **verbatim**, typos and slang included, into the raw list with an
   evidence line:

   ```
   "quote" [URL or source name · seen: YYYY-MM-DD · language · rating if any · opened | supplied]
   ```
3. Keep the original language. If a translation helps, put it on the next line,
   a line starting `gloss:`. Never use the gloss as the quote.
4. Take the full range: one-star and five-star reviews, complaints about competitors, questions asked
   before buying, reasons given for returns. Positive quotes show desires; negative ones show pains and
   objections.

Support messages, public reviews and your own customer records may contain personal data. Before a quote
enters the list, strip names, usernames and handles, emails, phone numbers, addresses, and order, ticket or
account numbers. For a private message, the source in the evidence line is the channel (for example "support
inbox"), never the sender or the ticket. If a public page address contains a username or profile id,
cite the page or thread without it, or just the site and the date. For a post on X, cite
`x.com/i/status/<post id>`, which leaves out the author's handle. Keep only the words about the problem.

## 2. Cluster

Group quotes by what they are about. Each cluster gets one type:

| Type | It answers |
|---|---|
| PAIN | what goes wrong, what hurts, what it costs them |
| DESIRE | the result they want, in their words |
| OBJECTION | why they hesitate or do not buy |
| TRIGGER | the moment or event that made them look for a solution |
| COMPARISON | what they measure you against: another product, a different kind of solution, their current habit |

Report each cluster as:

```
## <TYPE>: <cluster name>: N mentions from M sources
1. "quote" [evidence line]
2. "quote" [evidence line]
...
Most frequent phrasing: #n
Most vivid phrasing: #n
```

Rules:

- **N equals the number of numbered quotes listed.** It is a count, not an estimate of how common the
  problem is.
- **M counts distinct sources**, so one angry thread does not look like a market-wide pain.
- Pick the most frequent phrasing as representative. Mark the most vivid one separately; it is often an
  outlier.
- A quote that fits two clusters goes into both, and each cluster notes the overlap.

## 3. Rank

Order clusters by N, then by M. Flag clusters with one source only as `SINGLE SOURCE`. Flag clusters whose
newest quote is older than the freshness window in `CONTEXT.md` as `STALE`.

## 4. Turn clusters into angles

Only a cluster that meets the "Minimum before an angle" in `CONTEXT.md` (mentions and distinct sources)
becomes an angle card. If the minimum is `no data` or blank, ask the user for it. If they do not know,
use at least 3 quotes from at least 2 sources and say so in the output: fewer than three quotes is an
anecdote, and a single source can be one loud thread. List the clusters below the minimum by name so the
reader sees what was held back.
For each cluster that meets it, write:

```
ANGLE: <one line, in the customer's words where possible>
From cluster: <name> (N / M)
Awareness stage: unaware | problem-aware | solution-aware | product-aware | most aware
Message candidates:
- <a hook or headline using a verbatim phrase, with the quote number it comes from>
- <a second one from a different phrasing>
What we would need to prove it: <the evidence a sceptical reader asks for>
```

Match the message to the stage. The five stages follow Eugene Schwartz's levels of awareness
(Breakthrough Advertising): unaware readers need the problem named; problem-aware readers need to learn a
solution exists; solution-aware readers need to see why this approach; product-aware readers need a
reason to pick you over what they know; the most aware need the offer. A message aimed at the wrong stage
is wasted.

Every message candidate points back to a numbered quote. A candidate with no quote behind it is marked
`OUR WORDS, NOT THEIRS`.

## 5. Output

Write to the cluster bank from `CONTEXT.md` (template in `cluster-bank.template.md`):

1. The ranked cluster list with counts.
2. The angle cards.
3. A "not found" section: questions you went looking for and found no customer language on, and sources
   that could not be read.

## Self-check before replying

- Is every quote verbatim, in its original language, with an evidence line?
- Does every N match the quotes listed under it?
- Are personal details stripped from every quote and from every evidence line, including usernames or
  profile ids in page addresses?
- Does every message candidate point to a quote number, or say it is our words?
- Is the "not found" section present?

## Failure modes

| Symptom | Cause | Fix |
|---|---|---|
| Clusters read like a marketing brief | quotes paraphrased or summarised | copy verbatim; paraphrase only in the cluster name |
| A pain looks huge but comes from one thread | counted mentions, ignored sources | report N and M; flag single-source clusters |
| Polished "customer language" that nobody said | translations used as quotes | original language only; translations on a labelled line |
| Copy built on the most dramatic review | vivid quote taken as representative | representative = most frequent; vivid marked separately |
| Personal data in the bank | messages pasted raw, or a username in a review link | strip names, handles, emails, phones, addresses and order, ticket or account numbers first; cite review pages without usernames or profile ids |
