Voice-of-customer mining: customer language as evidence

The agent collects what customers really say, counts it honestly, and only then builds messaging angles from it.

Research, intermediate. Published

What it does

The agent reads reviews, forums, comments and customer messages and copies each statement verbatim, typos included, in its original language and with an evidence line. It strips personal data from messages. Quotes are grouped into clusters of five types: pain, desire, objection, trigger and comparison. Each cluster shows the number of quotes and the number of distinct sources, and marks the most frequent and the most vivid phrasing separately. The strongest clusters become angle cards with one of Eugene Schwartz's five awareness stages and headline candidates. Every candidate points to the quote number it comes from.

When to use it

When to use it

  • Before writing a product page, an ad or an email.
  • You are planning content and want topics from questions customers actually ask.
  • You are positioning an offer and want to know what customers compare you with.
  • You have plenty of reviews and messages but nobody has read them systematically.

When not to use it

  • You have no customers yet and no public conversation about the problem. Talk to a few people first.
  • You need numbers on market size. This skill counts quotes; it does not measure demand.
  • The only source is personal data you may not process for this purpose.

Decision table

SituationWhat the skill does
The source opensCopies quotes verbatim with an evidence line
The source needs a login or an exportGives you a link, says what to copy, labels the result supplied
A quote in another languageKeeps the original and adds an explanation on its own gloss line
A message or review with personal dataStrips name, username, email, phone, address and order, ticket or account numbers before saving
A cluster from a single threadFlags it as single source
A headline candidate with no quote behind itLabels it as our words, not theirs

Template

The method: collecting verbatim quotes with sources, pain, desire and objection clusters with honest counts, ranking, and angle cards.

SKILL.mdDownloadSKILL.md
---
name: voice-of-customer-mining
description: Collect what customers actually say about a problem or product category from reviews, forums, comments and support messages, keep every quote verbatim with its source, cluster the quotes by pain, desire and objection with honest counts, and turn the clusters into angles and message candidates. Use it before writing copy, planning content, positioning an offer, or when someone asks "what are our customers really saying".
---

# Voice-of-customer mining

Copy written from a summary ("customers care about quality") sounds like every competitor's copy. The
words that land are the ones customers already use: specific, sometimes rough, often misspelled. This
workflow collects those words as evidence, counts them honestly, and only then turns them into angles. It
never invents a quote and never passes off a translation or a paraphrase as the customer's voice.

Read `CONTEXT.md` next to this file first. It names the product or problem, the sources you may mine, the
competitors whose customers you may read, the language that counts as primary, and where the cluster bank
lives.

If `CONTEXT.md` is missing, or a field still holds `___`, ask the user for it before collecting anything,
in the language they write in. Ask for the competitor list and the sources by name, for example:
„Dodaj swoją listę konkurentów i źródeł opinii: …" or "Add your competitor list and the places your
customers talk: …". Never make up a competitor or source list yourself. "No competitors" is a valid
answer; write it down and carry on.

## Before you start / What you need

- **An AI agent that can open web pages**: Claude Code (https://code.claude.com/docs/en/overview) has web
  search and page fetch built in; Codex CLI (https://developers.openai.com/codex/cli) and Cursor
  (https://cursor.com/docs) work too. Without page access the user pastes the content and every quote is
  `supplied`.
- **Optional: Playwright MCP** (Apache-2.0, https://github.com/microsoft/playwright-mcp) for review pages
  that only load in a real browser. Needs Node.js. Claude Code: `claude mcp add playwright npx
  @playwright/mcp@latest`; Codex: `codex mcp add playwright npx "@playwright/mcp@latest"`. Do not sign in
  anywhere with it; a page behind a login stays `supplied`.
- **Optional: the official X API** (https://docs.x.com/x-api/introduction) to read public posts. The user
  creates an app in the Developer Console (https://console.x.com), buys pay-per-use credits, sets a
  spending limit, and generates an app-only token. The token lives in an environment variable in the
  user's shell profile or a `.env` file listed in `.gitignore`, never in the chat or in `CONTEXT.md`. Use
  the recent search endpoint (https://docs.x.com/x-api/posts/search/introduction).
- **Optional: review exports** from the user's own store or review app, stripped of personal data first.

If a tool is missing, say so once and mark those sources `supplied`.

## 1. Collect

For each source in `CONTEXT.md`:

1. Open it and read what is actually there. If it will not load or needs a login, do not guess: give the
   reader a ready link and say exactly what to copy back, then work from what they paste.
2. Copy each relevant statement **verbatim**, typos and slang included, into the raw list with an
   evidence line:

   ```
   "quote" [URL or source name · seen: YYYY-MM-DD · language · rating if any · opened | supplied]
   ```
3. Keep the original language. If a translation helps, put it on the next line,
   a line starting `gloss:`. Never use the gloss as the quote.
4. Take the full range: one-star and five-star reviews, complaints about competitors, questions asked
   before buying, reasons given for returns. Positive quotes show desires; negative ones show pains and
   objections.

Support messages, public reviews and your own customer records may contain personal data. Before a quote
enters the list, strip names, usernames and handles, emails, phone numbers, addresses, and order, ticket or
account numbers. For a private message, the source in the evidence line is the channel (for example "support
inbox"), never the sender or the ticket. If a public page address contains a username or profile id,
cite the page or thread without it, or just the site and the date. For a post on X, cite
`x.com/i/status/<post id>`, which leaves out the author's handle. Keep only the words about the problem.

## 2. Cluster

Group quotes by what they are about. Each cluster gets one type:

| Type | It answers |
|---|---|
| PAIN | what goes wrong, what hurts, what it costs them |
| DESIRE | the result they want, in their words |
| OBJECTION | why they hesitate or do not buy |
| TRIGGER | the moment or event that made them look for a solution |
| COMPARISON | what they measure you against: another product, a different kind of solution, their current habit |

Report each cluster as:

```
## <TYPE>: <cluster name>: N mentions from M sources
1. "quote" [evidence line]
2. "quote" [evidence line]
...
Most frequent phrasing: #n
Most vivid phrasing: #n
```

Rules:

- **N equals the number of numbered quotes listed.** It is a count, not an estimate of how common the
  problem is.
- **M counts distinct sources**, so one angry thread does not look like a market-wide pain.
- Pick the most frequent phrasing as representative. Mark the most vivid one separately; it is often an
  outlier.
- A quote that fits two clusters goes into both, and each cluster notes the overlap.

## 3. Rank

Order clusters by N, then by M. Flag clusters with one source only as `SINGLE SOURCE`. Flag clusters whose
newest quote is older than the freshness window in `CONTEXT.md` as `STALE`.

## 4. Turn clusters into angles

Only a cluster that meets the "Minimum before an angle" in `CONTEXT.md` (mentions and distinct sources)
becomes an angle card. If the minimum is `no data` or blank, ask the user for it. If they do not know,
use at least 3 quotes from at least 2 sources and say so in the output: fewer than three quotes is an
anecdote, and a single source can be one loud thread. List the clusters below the minimum by name so the
reader sees what was held back.
For each cluster that meets it, write:

```
ANGLE: <one line, in the customer's words where possible>
From cluster: <name> (N / M)
Awareness stage: unaware | problem-aware | solution-aware | product-aware | most aware
Message candidates:
- <a hook or headline using a verbatim phrase, with the quote number it comes from>
- <a second one from a different phrasing>
What we would need to prove it: <the evidence a sceptical reader asks for>
```

Match the message to the stage. The five stages follow Eugene Schwartz's levels of awareness
(Breakthrough Advertising): unaware readers need the problem named; problem-aware readers need to learn a
solution exists; solution-aware readers need to see why this approach; product-aware readers need a
reason to pick you over what they know; the most aware need the offer. A message aimed at the wrong stage
is wasted.

Every message candidate points back to a numbered quote. A candidate with no quote behind it is marked
`OUR WORDS, NOT THEIRS`.

## 5. Output

Write to the cluster bank from `CONTEXT.md` (template in `cluster-bank.template.md`):

1. The ranked cluster list with counts.
2. The angle cards.
3. A "not found" section: questions you went looking for and found no customer language on, and sources
   that could not be read.

## Self-check before replying

- Is every quote verbatim, in its original language, with an evidence line?
- Does every N match the quotes listed under it?
- Are personal details stripped from every quote and from every evidence line, including usernames or
  profile ids in page addresses?
- Does every message candidate point to a quote number, or say it is our words?
- Is the "not found" section present?

## Failure modes

| Symptom | Cause | Fix |
|---|---|---|
| Clusters read like a marketing brief | quotes paraphrased or summarised | copy verbatim; paraphrase only in the cluster name |
| A pain looks huge but comes from one thread | counted mentions, ignored sources | report N and M; flag single-source clusters |
| Polished "customer language" that nobody said | translations used as quotes | original language only; translations on a labelled line |
| Copy built on the most dramatic review | vivid quote taken as representative | representative = most frequent; vivid marked separately |
| Personal data in the bank | messages pasted raw, or a username in a review link | strip names, handles, emails, phones, addresses and order, ticket or account numbers first; cite review pages without usernames or profile ids |

Other files

CONTEXT.template.mdYour settings: topic and decision, sources with their mode, primary language, freshness window, the threshold for an angle and the cluster bank path.
CONTEXT.template.mdDownloadCONTEXT.template.md
# Voice-of-customer mining: your context

Copy this file to CONTEXT.md next to SKILL.md (Claude Code), or next to your AGENTS.md or rules file
(Codex, Cursor), and fill in every blank. Write `no data` for anything you do
not know yet. Never put passwords, access keys, card numbers or customer personal data here.

## What we are listening for

(The product, category or problem, and the decision the result feeds.)
Topic: ___
Decision it feeds: ___
Example (not your data): reusable food wraps; feeds the headline and first section of the product page.

## Sources

(Where customers talk about this. Mark each opened if the agent can open it, supplied if you will copy it in.)
| Source | Why it matters | Mode |
|---|---|---|
| ___ | ___ | opened / supplied |
Example (not your data): reviews on our own product pages, supplied from an export; competitor product reviews, opened where pages load; two hobby forums, opened; our support inbox, supplied after stripping personal data.

## Competitors

(Add your own competitor list: the products or brands whose customers' words you want read. The agent asks
for this list if it is empty and never makes one up. Write `none` if you only want your own customers.)
| Competitor | Where their customers talk |
|---|---|
| ___ | ___ |
Example (not your data): two rival brands of the same product; their product review pages and one comparison thread.

## Primary language

(The language of quotes that count as your customers' voice. Quotes in other languages are hints.)
Primary language: ______
Example (not your data): the local language of the market we sell in.

## Freshness window

(After how long a cluster with no newer quote is marked STALE.)
Freshness window: ______
Example (not your data): 12 months.

## Minimum before an angle

(How many mentions and sources a cluster needs before it becomes an angle card. If you write `no data`,
the agent asks, and without an answer uses at least 3 quotes from at least 2 sources.)
Minimum before an angle: ______
Example (not your data): at least five quotes from at least three sources.

## Cluster bank

(Where the output file lives.)
Path: ___
Example (not your data): ./research/voice-of-customer.md
cluster-bank.template.mdThe output file template: ranked clusters, clusters with numbered quotes, angle cards and a not-found section.
cluster-bank.template.mdDownloadcluster-bank.template.md
# Voice of customer: (topic)

Last updated: YYYY-MM-DD
Sources read: (list, each with opened or supplied)

## Ranked clusters

| Rank | Type | Cluster | N | M | Flags |
|---|---|---|---|---|---|
| 1 | PAIN | (name) | (n) | (m) | (SINGLE SOURCE, STALE, or none) |

## PAIN: (cluster name): N mentions from M sources
1. "(verbatim quote)" [source · seen: YYYY-MM-DD · language · opened | supplied]
2. "(verbatim quote)" [source · seen: ...]
Most frequent phrasing: #(n)
Most vivid phrasing: #(n)

## Angle cards

ANGLE: (one line in the customer's words)
From cluster: (name) (N / M)
Awareness stage: (unaware | problem-aware | solution-aware | product-aware | most aware)
Message candidates:
- (hook using quote #n)
- (second hook using quote #n)
What we would need to prove it: (evidence)

## Not found

- (question we looked for and found no customer language on)
- (source that could not be read, and why)

What you need

Install

  1. Claude Code: copy the folder to .claude/skills/voice-of-customer-mining/ in your project, or to ~/.claude/skills/voice-of-customer-mining/ for every project. Claude loads the skill when its description matches the task, or when you type /voice-of-customer-mining.
  2. Codex: paste the contents of SKILL.md into AGENTS.md at the repository root, or into the global ~/.codex/AGENTS.md.
  3. Cursor: add SKILL.md as a project rule, or paste it into AGENTS.md.
  4. A chat with no files: paste SKILL.md at the start of the conversation, then paste opinions in batches.
  5. Copy CONTEXT.template.md to CONTEXT.md next to SKILL.md (Claude Code), or next to your AGENTS.md or rules file (Codex, Cursor), and fill in every field.
  6. Create the cluster bank from cluster-bank.template.md at the path set in CONTEXT.md.

It's working if

  • Every quote in the bank has a source, a date, a language and how it was obtained.
  • The count on each cluster matches the number of numbered quotes under it.
  • Clusters carry a type: pain, desire, objection, trigger or comparison.
  • Every headline candidate points to a quote number or is labelled as our words.
  • The bank holds no names, usernames, email addresses, phone numbers, or order, ticket or account numbers.
  • The file ends with a not-found section.

Requirements

  • Access to opinion sources: public pages, a review export or pasted messages.
  • An agent that can open web pages. Without one, you paste the content yourself.
  • A lawful basis for processing customer messages, if you use them.

Questions

Why count sources separately from quotes?

One long thread can produce a dozen quotes about the same thing. Without the source count it looks like a market-wide problem. Two numbers side by side show whether it is a common pain or one loud conversation.

Can I translate quotes into my site's language?

To understand them, yes, on a separate labelled line. As quotes, no. A translated quote reads smoothly, but nobody said it that way, and the whole value of this method is the words customers really use.

How is this different from sentiment analysis?

Sentiment analysis tells you whether opinions are good or bad. This skill pulls out specific phrasings, groups them by what they are about and turns them into messaging angles with a quote attached. The result feeds copywriting directly.

Where it fits

This skill applies the research agent template's evidence discipline to one kind of material: customer words. The angle cards feed whatever copy or content brief you write next. In an online store it pairs with the store agent instructions template, where product descriptions draw on the cluster bank.

All skills

Sources

Questions about setting these up go in the Discord.Join free