Synthetic users: where AI research personas break, and where they earn a place
Can AI personas replace user testing? No as evidence, yes as preparation. A placement rule for where synthetic users help and where real interviews matter.
No, not as evidence. Yes, as preparation. A synthetic user, a model prompted to answer as if it were a member of your audience, is useful for rehearsing a discussion guide, pressure testing a set of questions, and generating hypotheses worth checking against real people. It is not useful as a stand in for the people themselves, because it cannot surprise you, cannot contradict your assumptions from lived experience, and gets least reliable exactly where the audience sits furthest from what the model has seen before.
That placement rule, before the study to prepare, never after the study as evidence, is the whole argument. What follows is why it holds.
The pitch
Synthetic users are being sold as a replacement for recruiting real participants. The pitch is speed: no scheduling, no incentive budget, no waiting a week for five sessions to land on the calendar. Ask a model to answer as if it were a member of your audience, and you get a transcript back in seconds. For a team under deadline pressure, that is a genuinely tempting trade.
The risk is what it looks like when it goes wrong. It does not look like an obvious failure. It looks like a confident, well written research summary about people who do not exist, sitting in a deck next to real findings with no visible seam between them.
What a synthetic user can honestly do
Where a synthetic answer belongs
Preparation
Rehearse the guide, widen the question list, pressure test the wording.
Hypothesis
Worth checking against real people. Labelled as a hypothesis in the report, never as a finding.
Evidence
Comes from real participants only.
Before the studyAfter the study
Used before a study, a synthetic user is closer to a rehearsal partner than a participant. Run a draft discussion guide past one and it will surface a leading question, a term your real participants might not use, or a gap where you have three questions about the happy path and none about what happens when something breaks. It will generate hypotheses: guesses about what a user might struggle with, worth carrying into the real sessions to check. None of that requires the synthetic answer to be true. It only has to be plausible enough to stress test the instrument, the way a colleague playing devil's advocate helps you find the weak questions in a plan before it ships.
That is genuine value, and it is available for free before a single real session is scheduled.
What it cannot do
A synthetic user cannot surprise you. It can only rearrange what you already told it.
The limit that does not go away with a better prompt
A model answering as a persona is drawing on patterns in its training data, shaped by the prompt you wrote. It cannot contradict your assumptions from a lived experience it does not have, and it cannot report an emotion, a hesitation, or a detail that surprises the person asking the question, because there is no experience underneath the answer to surprise anyone with. Ask it what a user might say, and it will produce something reasonable. Ask a real person, and you sometimes get an answer that reorganizes the whole project.
Plausible, generic, unsurprising
- Reflects patterns already in the prompt and the training data.
- Reads as reasonable because plausible is what it optimizes for.
Specific, sometimes contradictory, occasionally the whole finding
- Carries a detail nobody would have thought to invent.
- Can echo something said about an entirely different product, months apart.
The gap is widest exactly where it matters most: audiences far from what a general purpose model has seen a lot of. A synthetic user standing in for a highly specific or underrepresented group is not drawing on a rich internal picture of that group. It is drawing on whatever thin, generic signal exists about them, and it will not tell you that is what it is doing. The confidence of the answer does not shrink to match the thinness of what it is based on.
Two moments no model would have produced
I ran the physician interviews behind Avalon, the clinical mobile EMR I designed at CureMD, myself. When we tested an early direction for the Provider Note, one that captured the full clinical record in a single long entry pass, physicians described it the way they had separately described the old app they were replacing: something to get through, not something that helped them care for the patient in front of them. That specific echo, the same resignation surfacing for two different products a year apart, was the finding that reframed the whole feature. No prompt produces that. It depended on a person remembering a sentence from a different conversation and recognizing the same feeling in a new one.
SeniorConnect USA makes the same point from a different direction. The reader research there set constraints before a single screen existed: low vision, unsteady taps, jargon aversion, phone first use, bilingual households. Readers over fifty choosing internet service are not a group most general purpose models have deep, specific signal about, and a synthetic version of that reader would most likely default to a generic older adult stereotype rather than the actual, specific constraints research surfaced. The constraints came from real people describing their own experience, not from a plausible guess at what an older reader might say.
The placement rule
Put the two together and the rule is simple to state and easy to violate under deadline pressure: synthetic before the study, to prepare. Never after the study, as a substitute for the study. Anything a synthetic user produces is a hypothesis, and the report says so in those words, not folded in as if it came from someone real.
This is where a real research and usability work engagement earns its cost over the tempting shortcut. The fast version, skipping straight from a synthetic pass to a set of recommendations, produces a deck that reads well and rests on nothing. The honest version uses the synthetic pass to sharpen the questions, then still goes and asks real people, especially when the audience sits outside what a general model has seen a lot of.
Takeaways
Treat a synthetic user as a rehearsal tool. Use one to pressure test a discussion guide and generate hypotheses worth checking, and stop there. Never let a synthetic answer stand in for a real quote in a report, because the confidence of the answer does not track the thinness of what it is based on. The gap is widest for audiences a general model has thin signal about, which is often exactly the audience a project most needs to get right. And the finding that actually moves a project, an unprompted echo across two conversations, a constraint only a real reader would name, tends to come from the session a synthetic pass would have skipped.
Can AI personas replace user testing?
Not as evidence. As preparation, to rehearse a guide or generate hypotheses to check, yes.
When is a synthetic user actually useful?
Before a study, to pressure test the discussion guide and widen the question list. Anything it produces should be treated as a hypothesis, not a finding.
Why does the gap get worse for specific or underrepresented audiences?
A model draws on patterns in its training data. An audience with thin representation there gets a thin, generic answer back, and the answer does not announce its own thinness.
This is the other half of a question I have written about before: using LLMs in qualitative research synthesis covers what a model can honestly do once real interviews already exist. This post is about the step before that, deciding whether the interviews need to be real in the first place.