Does Human-in-the-Loop AI Annotation Really Work? 2025 Study
Qualitati Research Team · 2026-06-28 · 7 min read
Last updated: June 28, 2026
Short answer
Putting a human "in the loop" does not automatically fix AI annotation. In a pre-registered 2025 experiment with 410 annotators and over 7,000 annotations, Schroeder, Roy, and Kabbara (MIT) found that showing annotators an LLM's suggested label anchored their judgment, shifted the overall label distribution, and made people more confident — without making them faster. Critically, when those AI-assisted labels were then used to score the same AI, the model's measured performance was inflated. Human review is necessary, but it is not a guarantee of validity.
Key takeaways
- According to Schroeder et al. (2025), LLM label suggestions caused a strong anchoring effect: annotators "significantly" changed their label distribution toward the model's suggestions.
- LLM assistance increased annotator self-confidence but did not make annotation faster.
- Using AI-assisted labels as the "gold" set to evaluate the AI inflated the model's reported performance — a circularity trap.
- The risk is highest for subjective tasks (sentiment, stance, framing, theme judgments) where there is no single correct answer — exactly the tasks qualitative researchers care about.
- "Human-in-the-loop" is a workflow, not a validity claim. How you sequence AI and human judgment changes the result.
What did the study test?
The researchers ran a pre-registered online experiment on subjective text-annotation tasks — the kind where reasonable coders disagree. They recruited 410 crowdworker annotators who produced more than 7,000 annotations across conditions that varied whether, and how, an LLM suggestion was shown alongside each item. Because the tasks were subjective, there was no single ground-truth label, so the team measured how the distribution of human labels moved when an AI suggestion was present versus a baseline with no suggestion.
How much does the AI suggestion sway human coders?
A lot. According to Schroeder, Roy, and Kabbara (2025), annotators "strongly took the LLM suggestions, significantly changing the label distribution compared to the baseline." In other words, the same person shown an AI-proposed label tends to ratify it rather than independently re-judge the item. This is consistent with the broader literature on automation bias and anchoring: when a confident-looking default is on screen, people drift toward it.
Two secondary effects matter for research operations. First, annotators reported higher self-confidence when given LLM suggestions — they felt more sure even as their independence eroded. Second, the suggestions did not speed annotation up, so the common justification ("a human glance is fast and cheap") loses one of its main selling points.
Why does this inflate AI accuracy?
Here is the circularity the paper highlights. Suppose you generate labels with a human-plus-LLM workflow, then use those same labels as the benchmark to measure how good the LLM is. Because the human labels have already been pulled toward the LLM's outputs, the model is effectively being graded against an answer key it helped write. The study found that "reported model performance significantly increases" under this setup. Any downstream analysis built on those labels — an effect size, a prevalence estimate, a theme frequency — inherits the same distortion.
AI-first vs human-first annotation: a comparison
| Dimension | AI-suggests-first (review mode) | Human-first (blind, then reconcile) |
| Anchoring risk | High — label distribution shifts toward the model | Low — human judgment recorded before seeing AI |
| Speed | No measured gain (Schroeder et al., 2025) | Slower up front, but auditable |
| Confidence calibration | Inflated self-confidence | Truer to actual disagreement |
| Safe to evaluate the AI? | No — circular, inflates metrics | Yes — independent reference labels |
| Best use | High-volume, low-stakes triage | Reliability checks, validity-critical coding |
What this means for qualitative researchers
Qualitative coding is a subjective annotation task by another name. If your codebook involves interpretation — sentiment, motivation, stance, or theme assignment — the same anchoring dynamics apply when an AI proposes codes and a human "approves" them. Three practical implications:
- Don't measure AI agreement with AI-assisted labels. To estimate how well an AI coder matches humans, collect at least some blind human codes first, then compare. This keeps your inter-rater reliability estimate honest.
- Separate generation from validation. The coders who reconcile against the AI should not be the source of your reliability statistics. Reserve an independent, AI-naive sample as your reference.
- Treat "human-in-the-loop" as a design choice, not a stamp of approval. Document the order of operations: did the human see the AI label before or after deciding?
This is why tools that support qualitative analysis should make the workflow explicit. In Qualitati's QDA Workspace and ThemeLens thematic-analysis pipeline, AI-assisted coding is designed to keep an audit trail and to let researchers reconcile AI suggestions against their own judgment rather than silently inheriting them — so that reliability can be checked, not assumed.
Limitations and open questions
The study used crowdworkers on short subjective tasks, not trained qualitative researchers working a deep codebook, so the size of the anchoring effect among experts is an open question — experts may resist more, or may anchor just as readily under time pressure. The work is a 2025 arXiv preprint; readers should confirm its status when citing. And "blind-first" annotation costs time and money, so the right balance between speed and independence depends on how validity-critical the labels are.
FAQ
Does human-in-the-loop AI annotation make labeling faster?
Not in this study. According to Schroeder et al. (2025), LLM suggestions did not make annotators faster, even though annotators felt more confident.
What is the anchoring effect in AI annotation?
It is the tendency for human annotators to shift their labels toward an AI's suggested label simply because it is shown. The 2025 study found this significantly changed the overall label distribution versus a no-suggestion baseline.
Can I use AI-assisted labels to evaluate my AI coder?
You should not. Because those labels are pulled toward the model's own outputs, they inflate the model's measured performance. Use independent, AI-naive reference labels for evaluation.
How do I keep AI-assisted qualitative coding valid?
Collect some blind human codes before showing the AI's suggestions, keep an audit trail, and separate the people who generate labels from the reference set used to measure reliability.
The bottom line
"Just put a human in the loop" is not a validity strategy on its own. The 2025 evidence is that AI suggestions anchor human judgment, raise false confidence, and — if you are not careful — quietly inflate the very metrics you use to trust the AI. For qualitative work, the fix is procedural: judge blind first, reconcile second, and never grade an AI against labels it helped produce.
Primary source: Schroeder, H., Roy, D., & Kabbara, J. (2025). Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks. arXiv:2507.15821.
This article is an independent editorial summary of third-party research. Figures and findings are attributed to the cited authors; Qualitati is not affiliated with them. Last updated: June 28, 2026.