← All writing

Using LLMs in qualitative research synthesis without losing the user

Can AI do thematic analysis of user interviews? Yes, with conditions. A two pass method that keeps a verbatim quote behind every theme.

Yes, with conditions. A model can take a first pass at coding transcripts and clustering themes, and it will save you real time. What it cannot do is decide which cluster matters, notice what a participant did not say, or catch the moment two quotes only sound similar. That work stays with the researcher, and the method that works keeps it there on purpose: the model proposes, the researcher checks every proposal against the raw transcript, and nothing survives into the report without a verbatim quote behind it.

That rule, no theme without a quote, is the whole method. Everything below is what it takes to actually run it.

The cost that shrinks research

Interview synthesis is slow, and slow work gets cut first when a schedule tightens. Affinity mapping ten sessions by hand, printing transcripts, walking sticky notes across a wall, takes days that a project plan rarely has. So teams do the quiet thing: they run four interviews instead of ten, or they skip synthesis and go straight from a few conversations to a gut call about what they heard. Neither is a research failure exactly. It is a cost problem, and it has been a cost problem for as long as qualitative research has existed.

A model does not fix the cost problem by being smarter than a researcher. It fixes it by being fast at the part that was never the smart part: reading everything once, proposing a first grouping, and holding that grouping open for someone to argue with.

Where the help is real

Manual synthesis vs the two pass method

Transcribe

Raw sessions, nothing interpreted yet.

Code with quotes

Every proposed code carries the line it came from.

Check against source

Each cluster is walked back to the transcript.

Same start and end points as the manual path. The difference is where the evidence is attached.

A model is genuinely useful for a first pass at coding transcripts: reading every session and proposing candidate codes, the kind of granular labelling a human does but slower and less consistently across a long day of it. It is useful for clustering those codes into candidate themes, which is mechanical grouping work dressed up as insight work. And it is useful for surfacing contradictions across sessions, flagging that participant three said the opposite of participant seven on a point neither of them was asked about directly, something a tired researcher on session nine of ten can easily miss.

None of that is analysis. It is triage. It hands the researcher a shorter, better organized version of the same raw material, which is exactly the point at which the researcher's actual job starts.

Where it fails

Left alone, a model flattens what it summarizes. Ask it to describe how five participants felt about a broken flow and it will hand back a composite: calm, reasonable, generalized. It will not tell you that one participant went quiet for four seconds before answering, or that another one laughed in a way that meant something closer to giving up than to finding it funny. Tone, hesitation, the thing a participant circled back to twenty minutes later unprompted: none of that survives a text summary unless a person who was in the room, or who reads the full transcript, puts it back in.

The sharper failure is invention. Given a loose enough prompt, a model will propose a theme that sounds plausible and is not actually supported by anyone in the sample. Plausible is what it is optimizing for. Whether anyone actually said it was never part of the calculation. And it will miss absence entirely: the thing nobody mentioned, which is sometimes the finding. A model has no way to notice a silence it was never told to listen for.

The two pass method

The method that holds is two passes, with a rule that survives both.

Pass one. The model reads the full transcripts and proposes codes and clusters. Every proposed code is required to carry the line it came from: the actual quote, never a paraphrase. This step produces a rough map. The report comes later, and only after pass two.

Pass two. The researcher walks every cluster back against the raw transcript, one at a time. A cluster that holds up when you read the surrounding context stays. A cluster that only held up in isolation gets cut or rewritten. This is the slow part, and it is supposed to be, because this is the part where judgment actually happens.

No theme without a verbatim quote behind it.

The rule that survives every synthesis

This is where a research and interviews engagement earns its cost. Running the first pass fast is easy. Running the second pass honestly, cluster by cluster, against the real transcript, is the discipline a rushed project is most tempted to skip, and it is the part that makes the report defensible.

What manual synthesis caught

I did this by hand on the physician interviews behind Avalon, the clinical mobile EMR I designed at CureMD, before any of this tooling existed. The early version of the Provider Note captured the full clinical record in one long entry pass: history of present illness, severity, structured detail, mirroring exactly what the chart needed. In interviews, physicians described that long form the same way they had described the old app they were replacing: something to get through, not something that helped them care for the patient in front of them.

That specific echo, the same phrase surfacing for two different products a year apart, was the finding. It was not a theme a code-and-cluster pass would have generated on its own, because it depended on remembering how someone described a different piece of software months earlier and hearing the same resignation in a new sentence. A model summarizing that interview would likely have produced something reasonable and flatter: participants found the entry form lengthy. True, and useless, because it drops the exact detail that told us the form had repeated an old failure rather than introduced a new one.

That is the honest case for the two pass method rather than either extreme. Full manual synthesis catches connections like that one, but it does not scale past a handful of sessions before fatigue sets in. Handing the whole pass to a model scales, but it would have smoothed that connection into a generic summary and never told you it happened. The two pass method is the version where the model does the reading nobody wants to do twice, and the researcher stays the one who notices what an interview is actually echoing.

Takeaways

Use a model for the first pass: transcript coding, clustering, and flagging contradictions across sessions. Never let it produce the report on its own, because it will flatten tone and can invent a theme that sounds plausible and is not present. Require a verbatim quote behind every code from the start, so the second pass has something concrete to check. And walk every cluster back against the raw transcript yourself, because that is where the finding a summary would have smoothed over actually turns up.

The two pass method is not a shortcut around a slow process. It is a way to spend the time you do have on the half of synthesis that was always the point. The reading was never the hard part. The noticing was.

Can AI do thematic analysis of user interviews?

It can propose a first pass of codes and clusters. It cannot decide which cluster matters, and it should not be the last read of the transcript.

What is the biggest risk of using AI in qualitative synthesis?

Invented themes that sound plausible but no participant actually raised, and flattened detail that drops the tone or hesitation that made a moment significant.

Does this replace manual affinity mapping?

It replaces the first sort. The judgment still belongs to the researcher, who has to walk every cluster back against the raw transcript before it goes in a report.

If you have read this far, one related read is where I draw the same line for design work more broadly: AI belongs in the process, not the pixels, on where the acceleration is genuinely real and where it starts to erode trust.

ShareXLinkedIn
← PreviousYour prototype is an argument, not an artifactNext →Synthetic users: where AI research personas break, and where they earn a place

Keep reading