Skip to main content
← All founder guides
Four operational checks · 7 min read

Synthetic users: how to check the findings

Everyone in this field now says the same sensible thing: synthetic users are for hypotheses, not for decisions, and they complement real research rather than replacing it. That advice is correct and it stops one step early. Nobody says how to tell which is which.

You already have the findings. “Be careful” is not something you can do with them. Below are four checks that are.

Short answer

How do you know whether synthetic research findings are any good?

Run four checks on the findings themselves rather than on the tool: does it contain anything you did not put in, does it contain disagreement, does it describe real events or the idea of them, and has any of it been compared with answers from people who exist. Only the fourth produces a number.

The four checks

  1. Novelty: underline every claim that is not a recombination of your own brief. If nothing survives, the study reflected your beliefs back at speed
  2. Dissent: count positions, not respondents. Synthetic panels converge on the majority and on your question's framing, so unanimity means collapse rather than agreement
  3. Recall: a persona has no last time, so every account of the past is generated — look for the specific, slightly off-topic detail only a participant would produce
  4. Comparison: ask a handful of real people the same questions and keep the difference as a number. This is the only check that measures rather than judges
  5. The field agrees synthetic users suit desk research and hypothesis generation and skew shallow and favourable — the Nielsen Norman Group's guidance says exactly this
  6. The advice that stops at "complement, not replace" leaves the operational question open: these four checks are what closing it looks like

The four checks

1

The novelty check

Does the finding contain anything you did not put in?

Take your brief and the persona descriptions, then read the findings beside them. Underline every claim that is not a recombination of your own input. If nothing survives the underlining, the study told you what you already believed, at speed.

What failing it means: A study that only reflects your brief back is not wrong — it is empty. That is the most common failure and the hardest to notice, because agreement reads as confirmation.

2

The dissent check

Is there disagreement in it, or only an average?

Count the positions, not the respondents. Real groups scatter; synthetic panels converge towards the majority and towards the framing of the question. If every persona lands in the same place, the panel collapsed rather than agreed.

What failing it means: Convergence hides exactly the objection that matters — the rare one, from the segment least represented in the training data, which is also the one that kills launches.

3

The recall check

Does it describe an event, or the idea of an event?

Find every sentence about the past and ask whether it contains something only a participant would know: a figure, a workaround, a name, an inconvenient detail. Plausible accounts are smooth and general; recalled ones are specific and slightly off-topic.

What failing it means: This is the hard boundary. A persona has no last time, so any account of one is generated — in the same confident register as everything else, which is why it passes unnoticed.

4

The comparison check

Has any of it been put to people who exist, and what was the gap?

Take the three claims you would act on and ask a handful of real people the same questions. Then keep the difference as a number. Not a feeling that the panel was roughly right — the distance between what it predicted and what they said.

What failing it means: Almost nothing in this category offers it, because offering it means recording your own error rate. Without this check the other three are still judgement calls; with it, you have a measurement.

Why the fourth check is rare

The first three checks you can run yourself, today, on findings from any tool. The fourth needs something the tool has to be willing to do: put the same questions to real people and then keep the difference.

That is uncomfortable for a product, because the number it produces is an error rate. A tool with no comparison step can be confident indefinitely. A tool that runs one acquires a history of being wrong by specific amounts, and has to show it.

Here it works like this: the hypotheses from the synthetic stage become a survey with a public link, generated from your project rather than from a template. Nothing is scored until at least five real respondents have answered. Then the distance between what was predicted and what they said is computed and stored against the project.

And the honest part: that survey has to reach people, which is work that lands on you. A comparison with nobody in it is not a comparison — it is the same synthetic finding with a survey link attached.

Run the fourth check

A fifteen-minute interview produces the hypotheses, a synthetic panel rehearses the objections, and the same questions then go to real respondents so the two can be put side by side. The verdict is computed against thresholds rather than written by a model.

Start free →

15 min · No credit card · summary free

Frequently asked questions

What are synthetic users?+
AI-generated profiles that stand in for a user group and answer as if they were participants. They produce interview-shaped and survey-shaped output without anyone being interviewed or surveyed. The practical consequence is that a synthetic finding and a real finding look identical on the page, which is why the checks below are about the finding rather than about the tool.
Can synthetic users replace real user research?+
The settled answer across the field, including the Nielsen Norman Group's guidance, is no: they are suited to desk research and hypothesis generation, they tend towards shallow and favourable feedback, and they should complement rather than replace speaking with real people. The gap in that advice is that it stops at "be careful". Once you have the findings in hand you still have to decide what to do with them, and that requires checks rather than caution.
How accurate are synthetic users?+
They track real samples well on structured comparisons — ranking, sorting, reading direction on price — and not at all on anything that requires a past event, because no event occurred. So the accuracy question has no single answer: it depends entirely on whether the question you asked needs memory. A study that asked only comparative questions can be broadly right; the same panel asked what someone paid last year is generating fiction in the same tone of voice.
What is the strongest check?+
Ask a handful of real people the same questions and keep the difference. It is the only check that produces a number rather than a judgement, and it is the one almost no tool offers, because it requires the tool to be willing to record how wrong it was. Ours writes that gap down: the same hypotheses go out as a survey, nothing is scored until at least five real respondents have answered, and the distance between prediction and result is stored against the project.

More on this topic