Synthetic users: how to check the findings
Everyone in this field now says the same sensible thing: synthetic users are for hypotheses, not for decisions, and they complement real research rather than replacing it. That advice is correct and it stops one step early. Nobody says how to tell which is which.
You already have the findings. “Be careful” is not something you can do with them. Below are four checks that are.
Short answer
How do you know whether synthetic research findings are any good?
Run four checks on the findings themselves rather than on the tool: does it contain anything you did not put in, does it contain disagreement, does it describe real events or the idea of them, and has any of it been compared with answers from people who exist. Only the fourth produces a number.
The four checks
- Novelty: underline every claim that is not a recombination of your own brief. If nothing survives, the study reflected your beliefs back at speed
- Dissent: count positions, not respondents. Synthetic panels converge on the majority and on your question's framing, so unanimity means collapse rather than agreement
- Recall: a persona has no last time, so every account of the past is generated — look for the specific, slightly off-topic detail only a participant would produce
- Comparison: ask a handful of real people the same questions and keep the difference as a number. This is the only check that measures rather than judges
- The field agrees synthetic users suit desk research and hypothesis generation and skew shallow and favourable — the Nielsen Norman Group's guidance says exactly this
- The advice that stops at "complement, not replace" leaves the operational question open: these four checks are what closing it looks like
The four checks
The novelty check
Does the finding contain anything you did not put in?
Take your brief and the persona descriptions, then read the findings beside them. Underline every claim that is not a recombination of your own input. If nothing survives the underlining, the study told you what you already believed, at speed.
What failing it means: A study that only reflects your brief back is not wrong — it is empty. That is the most common failure and the hardest to notice, because agreement reads as confirmation.
The dissent check
Is there disagreement in it, or only an average?
Count the positions, not the respondents. Real groups scatter; synthetic panels converge towards the majority and towards the framing of the question. If every persona lands in the same place, the panel collapsed rather than agreed.
What failing it means: Convergence hides exactly the objection that matters — the rare one, from the segment least represented in the training data, which is also the one that kills launches.
The recall check
Does it describe an event, or the idea of an event?
Find every sentence about the past and ask whether it contains something only a participant would know: a figure, a workaround, a name, an inconvenient detail. Plausible accounts are smooth and general; recalled ones are specific and slightly off-topic.
What failing it means: This is the hard boundary. A persona has no last time, so any account of one is generated — in the same confident register as everything else, which is why it passes unnoticed.
The comparison check
Has any of it been put to people who exist, and what was the gap?
Take the three claims you would act on and ask a handful of real people the same questions. Then keep the difference as a number. Not a feeling that the panel was roughly right — the distance between what it predicted and what they said.
What failing it means: Almost nothing in this category offers it, because offering it means recording your own error rate. Without this check the other three are still judgement calls; with it, you have a measurement.
Why the fourth check is rare
The first three checks you can run yourself, today, on findings from any tool. The fourth needs something the tool has to be willing to do: put the same questions to real people and then keep the difference.
That is uncomfortable for a product, because the number it produces is an error rate. A tool with no comparison step can be confident indefinitely. A tool that runs one acquires a history of being wrong by specific amounts, and has to show it.
Here it works like this: the hypotheses from the synthetic stage become a survey with a public link, generated from your project rather than from a template. Nothing is scored until at least five real respondents have answered. Then the distance between what was predicted and what they said is computed and stored against the project.
And the honest part: that survey has to reach people, which is work that lands on you. A comparison with nobody in it is not a comparison — it is the same synthetic finding with a survey link attached.
Run the fourth check
A fifteen-minute interview produces the hypotheses, a synthetic panel rehearses the objections, and the same questions then go to real respondents so the two can be put side by side. The verdict is computed against thresholds rather than written by a model.
Start free →15 min · No credit card · summary free