A study on AI-moderated interviews blew up in my feed. I posted the headline hot take. Then researchers I respect did to my post what I should have done to the study: they read the method. Here's the tear-down I should have done in the first place, and the three lessons that survived it.
Last week I posted a hot take on a new Verasight study about AI-moderated interviews (Morris et al., 2026). It got about 1,700+ impressions and lit up with comments. And within a day, researchers I respect started doing to my post what I should have done to the study: reading the method.
Here's the humble pie. Every good researcher (analyst, statistician, and data scientist too!) knows the cardinal rule: before you trust a finding, interrogate the method and hunt for bias. Before you believe what any study says, you read the methodology: the study design, sampling strategy, and data processing. How the study was designed and the data gathered. It's the first move. And on a study I chose to amplify, I skimmed it. I led with the finding and let the method sit in the background.
Kyle Grant said it out loud, right in my comments: mine was the "second post I've seen repeating a completion rate headline without properly considering the study design." Thank you, Kyle; you were right. That's the first cardinal sin of research, and I committed it in front of my peers.
So here's the method read I should have done first. I owe everything below to the researchers who showed up in that thread.
Brief background on the study
So I have to first give Verasight credit: this is real work. They conducted a randomized experiment—3,160 people, the same four questions, and an AI-moderated interview against a written survey. They even flagged their own limits. The randomization is a real strength, which is why the study is worth reading closely instead of quoting quickly.
The three reported headlines from the report:
AI-moderated depth is real (compared to a survey). The AI arm produced ~5x more words and ~2x more usable evidence per person. About 80% of that lift came from adaptive follow-up probing.
Completion collapsed amongst AI interviews. 99.4% finished the written survey; 40.5% finished the AI interview. Most people dropped off at the hand-off to an external platform, before the interview started.
Dropout and completion weren't random. People who finished the AI-moderated interviews were more AI-optimistic and less representative. Weighting fixed only about a quarter of the gap.
Read the full Verasight report and findings here.
But this is only part of the story. The study design and methodological limitations require a heavy reframe to all three headlines.
The flaws & the fixes
Just to be clear: this study can't be read as a verdict on AI—at least not as is. Read it as a pilot, and it's a punch list. Five design choices change how you frame the findings, and almost all are fixable.
FLAW #1
The handoff: most people never reached the AI INTERVIEW
The 60-point completion gap is the number everyone repeats. But look at where it happened. Nobody was told they'd be redirected mid-survey to a live, AI-moderated video interview on a different platform—that caveat was buried in the write-up. Put yourself in the respondent's seat: you signed up for a five-minute survey, and suddenly you're asked to talk out loud, on camera, to an AI. People bail on that for mundane reasons—they're on transit, in the bath, sitting with family, or just not in the mood to perform.
The tell is in the study itself: of the people who actually reached the interview, 92% finished it. The drop happened before the AI said a word.
THE FIX: Tell people the exact format up front—spoken, on camera, roughly this long—and keep them inside one flow instead of bouncing them to a new site. Springing it on people also sits uneasily, in my read, with the EU AI Act's transparency principle—Article 50(1) expects people to be told when they're interacting with an AI system (EU AI Act, Art. 50(1)).
Flagged by Charles Allison and Kyle Grant.
FLAW #2
The burden: harder work, same pay
A spoken, probed interview is far more work than clicking radio buttons—more time, more thought, a microphone, a camera. The AI arm asked for all of that and paid the same incentive as the survey. Underpay the effort, and you don't lose a random slice of people; you lose everyone except the most eager and opinionated, who also tend to be the most tech-comfortable.
THE FIX: Match the incentive to the burden. Price a spoken interview closer to a micro-interview than a survey. Some of that 60% drop is a budget line.
FLAW #3
Depth by Default, But with Doubt
The AI arm produced roughly 5x the words. Some of that is unsurprising—people talk longer than they type, and the study pins about 80% of the lift on the follow-up probing, which a static survey text box simply can't do. That's not a finding; it's what the setup guaranteed.
But the size of the "win" is more dubious than a given. Both depth scores are measured on text. In the survey, the participant typed it. In the AI arm, a machine transcribed speech into text—and speech recognition rarely returns obvious gibberish. When the audio is unclear, it substitutes a plausible wrong word that reads fine on the page, and nothing flags it. Add that 16% of the AI completions were fraud-flagged, and that the quality rubric is scoring machine-transcribed text in the first place. Counting dropouts as zero, the raw premium is ~5x; measured on usable evidence rather than raw words, the honest premium is closer to 2x.
THE FIX: Score for usable evidence—concrete detail, specific example, causal explanation—not word count, and spot-check a sample of the raw audio against the transcript, especially for technical terms.
Flagged by Zeyu Si, who named the transcription-substitution risk directly.
FLAW #4
Limited scope & Sources
This tested one product on one panel, with one avatar script and one incentive structure. Change any of those, and the numbers move. It's a case study of a specific build. It would be unfair to extrapolate and apply the findings to any other AI-moderated tool, modality, and setup.
THE FIX: Read it as a single data point, and replicate it elsewhere before you generalize a whole method from one tool.
The fatal FLAW (#5)
COMPARING APPLES TO ORANGES
Now step back and look at what we just listed. The handoff, the burden, the mode, the probing—those weren't four independent problems. They were four faces of one. A clean experiment changes a single variable between two groups. This one changed at least five at the same time:
Modality: spoken voice vs. typed text.
Presence and social pressure: a dynamic video avatar vs. static form fields.
Interaction: real-time adaptive probing vs. single-shot prompts.
Platform: an external redirect from the survey to a separate interview platform vs. one native URL.
Technical hurdle: microphone and camera permissions vs. standard browser input.
Only one of those is "AI moderation." Move all five together, watch completion fall 60 points and depth rise 80%, and you still can't pin either number on the AI. The study set out to measure agentic AI. What it actually measured was unannounced, cross-site, modal friction.
To learn what AI moderation actually does, you hold everything else still and change one thing at a time. The two arms that would do the most work were never run:
The design, at a glance
What an honest version would test
| Arm | Mode | Moderator | Probing | In the study? |
|---|---|---|---|---|
| Control | Typed text | None | None | Yes |
| Treatment | Voice / video | AI avatar | Adaptive | Yes |
| Counterfactual 1 | Voice / video | Human | Adaptive | Missing |
| Counterfactual 2 | Typed text | AI chatbot | Adaptive | Missing |
Without Counterfactual 1, you can’t separate “AI” from “a human doing the same probing.” Without Counterfactual 2, you can’t separate “AI” from “talking instead of typing.”
THE FIX: Isolate the variables. Add targeted arms—a human voice interviewer running the same probes, an AI text chatbot embedded in the original survey, and an unmoderated audio recording with no probes. Just as important, strip out what shouldn't be variables at all: disclose the format up front, keep it on one platform, and pay for the added burden. Do that, and you can finally attribute the drop-off and the depth to specific causes.
One honest caveat: isolating the variables fixes internal validity—which lever moves which number. It doesn't fix the selection problem further down. No clean design changes who is willing to opt into an AI interview in the first place. You manage that one; you don't engineer it away.
The rest of the issues survive even a well-designed study.
So what can be applied? Not much.
Here's the takeaway I almost kept: AI interviews hand you a biased, unrepresentative sample. There's real signal for it. The people who finished skewed more AI-optimistic, more male, and more likely not to hold a degree. And when Verasight weighted the sample back toward the right demographics, it recovered only 12% of the attitudinal gap on one measure and 35% on another—roughly a quarter. Three-quarters of the skew survived a full demographic adjustment.
But apply the same discipline to that finding that I failed to apply to the headline. That quarter was measured under the same broken design—the ambush, the conflation, the mismatched expectations and pay. Drop-off and depth are tangled up in those artifacts. Fix them and the leftover selection might shrink toward the ordinary range every method carries. The study can't tell us how much because it never ran the clean version.
So "AI interviews give you unrepresentative samples" is exactly the kind of claim I shouldn't lift out of a confounded design.
What survives is narrower, and more useful—a principle, not a verdict on AI:
Every method suffers from some sort of self-selection bias. A survey, a human interview, an unmoderated diary study—someone always opts out. That alone isn't the problem.
Weighting only repairs the selection you measure: age, income, gender—the traits you capture. If you didn't measure it, you can't weight it. Here, that trait is willingness to talk to an AI, which plausibly tracks attitudes toward AI itself. That's why weighting recovered so little.
So the real risk with AI moderation is subtler: its bias can sit exactly where weighting can't reach—and whether that's a small, manageable cost or a fatal one, this study cannot say.
Brian Polk pushed the sharpest version in the comments: even with advance disclosure, the people who distrust AI may opt out on principle, so the pool stays skewed no matter how clean the mechanics get. He calls it the next iteration of the reproducibility crisis. He may be right—but notice that's a hypothesis this study can't settle either.
There's also a human echo of the same limit. Anna Khakho did an AI interview herself: the novelty carried the first half, then she realized the other side wasn't really listening, and she checked out. Rob Wallis named why it matters—real interviews run on rapport and trust, and automated probing can't build either. Dr. Susanne Friese offered the counterweight: the friction will likely soften as talking to AI becomes ordinary. Early days, not a dead end.
Isolation cleans up the design. It can't tell you who will show up, or how it feels to be interviewed by something that isn't really listening—and this study, tangled as it is, can't tell you either.
So what do you actually take from it?
Almost none of the findings—not cleanly. (Sad trombone) The depth is confounded, the drop-off is a handoff artifact, and you can't report the representation number without controlling for AI opt-in. If you came for a verdict on AI-moderated interviews, there isn't one here. That's not a failure of the study so much as a limit of what any single study trying to measure five things at once can tell you.
The value is the importance of method and study design. Read the right way, this study is a blueprint for the one you'd have to run yourself, and every flaw above is a variable to pin down before you trust a number:
Isolate the treatment—add a human-moderator arm and a text-based AI arm, so you know what "AI" is actually doing.
Disclose the format up front and keep it in one flow, if possible. Or, at the very least, allow participants to complete when possible.
Match the incentive to the burden.
Score usable evidence, not word count, and spot-check the audio against the transcript.
Target your sampling to backfill who drops, and don't expect weighting to rescue it.
That's the real deliverable: a checklist for finding out for yourself—on your own panel, your own tool, your own questions. The study can't rule on whether AI interviews are good or bad. You can, once you control for what it left tangled. It's the frame I'm using to design my own tests, and what I'll keep building into an AI-moderated-interview playbook.
None of this needed a better AI. It needed someone to read the method before repeating the headline—and my comments section did it for me. A HUGE thank you to Charles Allison, Kyle Grant, Kimmo Holm, Zeyu Si, Brian Polk, Rob Wallis, Anna Khakho, and Dr. Susanne Friese for sharing their brilliant minds with me.
What's the last research stat you shared before you read how it was built?

