AI for UX research in 2026: synthesis that doesn't hallucinate
How to structure notes, themes, and evidence so AI summaries are trustworthy, plus the tools that make the trail auditable.





The problem with AI synthesis was never that it produces summaries. It's that it produces confident summaries, and a confident summary of eleven interviews is indistinguishable from a real finding until someone challenges it in a meeting.
The fix is structural, not model-dependent. Build a pipeline where every claim can be traced to a quote, and hallucination stops being a risk you have to worry about. It becomes a thing you can see.
The rule: no claim without a link
Everything below follows from one rule. A synthesis output is only acceptable if each theme points to specific, timestamped source material.
Not "users found onboarding confusing." Instead: "users found onboarding confusing: P3 14:20, P7 08:45, P11 22:10." When a stakeholder pushes back, you play the clip. When the model invents a theme nobody said, the absence of a link exposes it immediately.
This one rule does more for trustworthy synthesis than any amount of prompt engineering.
Stage 1. Capture cleanly
Bad transcripts produce bad everything. Speaker separation matters more than word-level accuracy.
Granola is the current favourite for live sessions, because it takes your own sparse notes and enriches them against the transcript, which keeps you present in the interview rather than typing through it. Fireflies and Otter are the reliable workhorses, and Grain is strongest if you want shareable clipped moments as a first-class output.
For unmoderated work, Maze and UserTesting handle capture and structure together.
Stage 2. Tag before you summarise
This is the step people skip, and skipping it is why their AI summaries are mush.
Define a taxonomy first: a flat list of 10 to 20 tags that map to the questions you actually need answered. Then let AI apply the tags across the corpus, and spot-check maybe 20% of its assignments by hand.
The reason this order matters: a model asked to "find the themes" will find the themes that are linguistically salient, which correlates poorly with what's decision-relevant. A model asked to apply your taxonomy is doing a classification task it's genuinely reliable at.
Dovetail is the strongest tool for this because the tag-to-evidence link is native. Marvin and Looppanel are lighter alternatives that do the same job well for smaller teams.
Stage 3. Synthesise, adversarially
Now you can summarise. Three prompts worth running in sequence, every time:
"Summarise the evidence for each tag, citing participant and timestamp for every claim." Your baseline output.
"Which of these themes is supported by the fewest participants? Which could be an artefact of how I asked the question?" This surfaces the weak findings before someone else does. Models are surprisingly good at this when explicitly asked, and completely silent about it when not.
"What did participants say that contradicts the main finding?" The single most valuable question in the whole process. Synthesis naturally collapses toward consensus; this pulls the disconfirming evidence back out.
Stage 4. Write it up for humans
The output of research is a decision, not a document. Structure the write-up around what you want to change, with the evidence supporting it, and put the full themed corpus behind a link for anyone who wants to dig.
AI is good at the reformatting here: same findings as a one-paragraph summary, a slide, or a detailed appendix. It's poor at deciding what matters, which is the part you're paid for.
Where AI still shouldn't be trusted
Sample sizes below about eight. With five interviews you should read every transcript. The AI adds nothing and costs you the intuition you'd have built.
Sentiment scoring as a headline number. "72% positive sentiment" is a number that will be quoted in a deck for a year and means very little. Sentiment on a specific, narrow question is fine; sentiment as an aggregate metric is theatre.
Generating participant quotes. Obvious, but worth stating: if a quote isn't in a transcript, it doesn't go in the deck. Some tools will smooth a quote for readability. Turn that off.
Deciding what to build. Research tells you what's true. It doesn't tell you what to do about it, and a model that summarises the former will happily fabricate confidence about the latter.
The short version
Capture cleanly, tag deliberately, summarise adversarially, cite everything. Do that and AI synthesis is a real force multiplier, because you can run twice the research with the same rigour. Skip the structure and you've built a machine for producing plausible-sounding nonsense at speed.
More in the research & synthesis list.


