How to Report LLM Use in Qualitative Methods (2026)
Qualitati Research Team · 2026-09-14 · 10 min read
Last updated: September 14, 2026
Short answer
To report LLM use in qualitative research, state in the Methods section which model and version you used, when, through which interface, for which analytic step, with which prompts and settings, how many runs, and how humans checked the output. A June 16, 2026 scoping review found that 75% of LLM-assisted qualitative studies reported no parameter settings at all, so these details are where most papers fall short.
Knowing how to report LLM use in qualitative research is now a practical skill, not a formality. Journals ask for it. Reviewers check for it. And a methods section that just says "ChatGPT assisted with coding" gives a reader no way to judge your themes.
This guide summarizes what current publisher policies and the 2026 evidence say. It closes with a copy-ready checklist and a template paragraph.
Key takeaways
- A scoping review of 75 LLM-assisted qualitative studies published 2020–2025 found 74 named the model, but only 13 reported temperature and 56 (75%) reported no parameters at all (Kempny et al., 2026).
- No finished LLM-specific qualitative reporting guideline exists yet. COREQ+LLM is an extension of the 2007 COREQ checklist, and the EQUATOR Network still lists it as under development (Fehring et al., 2025).
- Nature Portfolio asks that LLM use be documented in the Methods section (Nature Protocols editorial policies). APA journals ask authors to describe AI use for coding or analysis in the method section (APA).
- Setting temperature to 0 does not make results reproducible, so report the number of runs and how much they varied.
- Use the LLM Methods Disclosure Checklist below before you submit.
Why reporting LLM use in qualitative research matters now
Reporting LLM use means giving readers enough detail to judge, and ideally repeat, what an AI system did to your qualitative data. In traditional qualitative work, credibility rests on a visible analytic path: who coded, how the codebook changed, and how themes were reached. When an LLM does part of that work, part of that path becomes invisible unless you document it.
The gap is measurable. In the June 16, 2026 scoping review in BMC Medical Research Methodology, Kempny, Frings, Rust, Meister, and Fehring examined 75 peer-reviewed empirical studies. OpenAI GPT models appeared in 93% of them. The review reports that:
- 61% of studies gave complete or partial prompts. The rest gave little or no prompting detail.
- 45% did not say how the model was accessed. Of the rest, 38 used an API or web interface and 5 ran the model locally.
- 97% included some human verification. But human–AI agreement ranged from 36% to 99%, so "we checked it" says little on its own.
The authors conclude that dedicated reporting guidelines are urgently needed. Until COREQ+LLM is published, researchers have to assemble the standard themselves.
What publishers currently ask for
Policies differ, but they converge on one point: AI used on research data belongs in the Methods, not only in the acknowledgments.
| Source (checked September 14, 2026) | Where to disclose | What it emphasizes |
| Nature Portfolio | Methods section, or a suitable alternative section | LLMs cannot be authors; use must be documented |
| APA Journals | Method section | How, when, and to what extent AI was used; keep prompts and outputs available |
| COREQ+LLM (in development) | Qualitative reporting checklist | Model selection, prompting, validation, bias mitigation |
Always read your target journal's own author guidelines. Many publish their own rules, and requirements change quickly.
Why "temperature 0" is not a reproducibility claim
Many methods sections say the analysis used temperature 0 "to ensure consistency." The evidence does not support that wording. In a June 2026 study of 690 API calls across 7 conditions, Tamba found that some borderline items still produced different verdicts under greedy decoding, and that top_p and top_k did not close the gap (Tamba, 2026). The paper recommends running several passes and reporting the variance, not a single output.
That matches an earlier result we covered: repeated identical analyses by LLMs often reach different conclusions. Some newer models also no longer accept a temperature setting at all. The honest report is therefore: settings used (or "not configurable"), number of runs, and how differences between runs were handled.
The LLM Methods Disclosure Checklist
This is a Qualitati-developed checklist. It combines the gaps in the Kempny et al. review with the publisher requirements above. It does not replace COREQ+LLM once that is published.
A. The model
- ☐ Provider, model name, and exact version or snapshot identifier
- ☐ Dates of use (models behind the same name are updated)
- ☐ Access route: web chat, API, a named research platform, or local deployment
B. The task
- ☐ Which analytic step the model performed: transcription, initial coding, codebook drafting, applying codes, theme synthesis, or quote retrieval
- ☐ The analytic approach it served: reflexive thematic analysis, codebook thematic analysis, or content analysis
- ☐ What was not delegated to the model
C. Inputs and settings
- ☐ Full prompts in an appendix or supplementary file, including system instructions
- ☐ Temperature, top_p, and output limits, or "not configurable" if the interface hides them
- ☐ How transcripts were split or batched to fit context limits
- ☐ De-identification applied before data reached the model
D. Stability and human review
- ☐ Number of runs per step, and how disagreements between runs were resolved
- ☐ Who reviewed the output, what share was checked, and the agreement statistic with its denominator
- ☐ Examples of codes or themes that humans changed or rejected
E. Governance
- ☐ Whether participants consented to AI processing of their data
- ☐ Data retention and training terms of the provider
- ☐ Researcher reflexivity: how using AI may have shaped interpretation
Template paragraph
"We used [model, version] from [provider], accessed via [interface] between [dates], to [task] for [n] de-identified transcripts. Prompts are provided in Supplementary File [X]. Settings were [temperature / top_p, or 'not user-configurable']. Each step was run [n] times; differences between runs were resolved by [procedure]. [Researcher role] reviewed [share] of AI-assigned codes (agreement: [statistic], n = [units]) and revised [description]. Theme development and interpretation were conducted by the research team. Participants consented to AI-assisted analysis."
Who this is for, and when not to use this approach
This is for academic qualitative researchers, UX researchers publishing case studies, and insights teams whose findings go to executives who ask "how did you get this?"
When not to use it as written: if an LLM only fixed grammar in your manuscript, most policies treat that differently from analysis. Check the journal's own copy-editing rules. And if your epistemology is strongly interpretivist, a checklist cannot stand in for a reflexive account of how the model shaped your reading of the data. Include that account as well.
Limitations and trade-offs
- Full prompts can be long. Put them in supplementary material rather than trimming them.
- Version identifiers can be unavailable in consumer chat interfaces. Say so. Silence reads as omission.
- Agreement statistics can mislead. High agreement on easy codes can hide disagreement on the interpretive ones. Report agreement per code where possible.
- Standards are still moving. This checklist reflects publicly available policies and evidence as of September 14, 2026. Treat it as a starting point that needs human review against your field's norms.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX, and customer insights teams. It runs AI-moderated interviews, focus groups, and conversational surveys, and analyzes the results. Its ThemeLens thematic analysis runs a documented map-reduce pipeline across up to 100 transcripts, maps codes to research questions, and anchors themes to participant quotes. That structure is easier to describe step by step in a methods section than an open-ended chat session. The QDA Workspace supports AI-assisted inductive and deductive coding with human review.
A platform does not write your disclosure for you. You still need to record dates, runs, and human review decisions yourself.
FAQ
Do I need to disclose AI use if I only used it for transcription?
Usually yes. Transcription creates the data you analyze, so name the tool and say how transcripts were checked. Confirm the requirement in your journal's author guidelines.
Is naming "ChatGPT" enough?
No. Kempny et al. (2026) found almost every study named the model, yet most omitted settings and deployment details. Give the version, dates, access route, prompts, and settings.
Where should prompts go?
Put a short description in the Methods and the full text in supplementary material. APA advises keeping prompts and outputs so editors can request them.
Does temperature 0 make my analysis reproducible?
No. A 2026 study found outputs could still differ at temperature 0. Report the number of runs and how differences were resolved.
Is there an official reporting guideline for LLMs in qualitative research?
Not yet. COREQ+LLM is in development. Until it is published, combine COREQ or SRQR with your journal's AI policy and a checklist like the one above.
Should I mention AI in the limitations section too?
Yes, if the model's role could have shaped your findings, for example by compressing themes or missing culturally specific meaning.
Conclusion
Good reporting of LLM use in qualitative research comes down to making the invisible steps visible: model, version, dates, prompts, settings, runs, and human review. The 2026 evidence shows settings and deployment are the most common omissions, so start there. If you want an analysis workflow that is easy to describe, start free with 30 credits or view transparent pricing.
This article is an independent editorial summary of publicly available policies and research as of September 14, 2026. It is not an official interpretation of any publisher's policy.