How to Run a Diary Study With AI (2026 Guide)
Qualitati Research Team · 2026-08-11 · 9 min read
Short answer. A diary study asks participants to log experiences in their own context over days or weeks, so you capture what actually happens instead of what people later remember. To run one with AI, keep the cadence low, make each entry short, let an AI moderator ask one contextual follow-up per entry, and reserve human judgment for sampling, consent, and interpretation. Compliance design matters more than tooling.
Last updated: August 11, 2026
Why run a diary study at all?
Most product teams already run interviews. A diary study is the method you add when the research question has a time dimension that a single sitting cannot reach: how a habit forms, when frustration peaks in a weekly workflow, how sentiment drifts after onboarding, or which moments in a real day the product never touches. The Nielsen Norman Group describes diary studies as a longitudinal method producing a self-reported record of behavior and attitudes that researchers analyze afterward to understand habits and patterns (NN/g).
The trade is real. Interviews are cheap to schedule and expensive to trust for frequency claims. Diary studies are the reverse: the data sits close to the moment, and the hard part is getting people to keep logging.
Key takeaways
- Diary studies reduce recall bias but introduce a compliance problem — plan for the compliance problem first.
- Published experience-sampling evidence puts average response rates near 78%, with predictable decay across days and across times of day.
- Sampling frequency is a design lever: spreading prompts across the day outperformed cramming them into a short burst in a 2026 Field Methods study.
- Onboarding, not the diary itself, is where most of the sample is lost — and lost non-randomly.
- AI helps most at two seams: one adaptive follow-up per entry, and cross-participant synthesis at the end. It does not fix a badly designed cadence.
What the compliance evidence actually says
Diary studies sit in the same family as the experience sampling method (ESM), which has been measured far more rigorously than UX practice usually acknowledges. Two findings are worth designing around.
1. Compliance is decent on average and decays predictably. In a meta-analysis of ESM studies, Rintala, Wampers, Myin-Germeys, and Viechtbauer (2019, Psychological Assessment) reported an average response rate of 78% (95% CI 74-82). Within that average, response rates were highest between 12:00 and 13:30 (83%) and lowest between 07:30 and 09:00 (56%), and participation declined across the study, reaching 73% by day 5 (PubMed).
Read that as three design instructions: do not prompt first thing in the morning, expect roughly a quarter of entries to be missing by the end of the first week, and treat "we got fewer entries on day 5" as a known property of the method rather than a study failure.
2. Frequency and onboarding decide your sample. Ohme, Charlton, Toth, Araujo, and de Vreese, publishing in Field Methods (online October 15, 2025; Volume 38, Issue 2, May 2026), randomized 250 Dutch participants across seven days into either a daily-intensive design (7 surveys spread across 14 hours) or an hourly-intensive burst design (12 surveys inside a 2-hour window). The spread-out condition achieved significantly higher compliance than the burst condition (Field Methods).
The attrition figure is the one to show stakeholders. Of 1,411 initial recruits, 250 — 17.82% — completed the mobile ESM protocol, with systematic dropout concentrated in onboarding. Participants who made it into the app skewed younger, more educated, and more tech-savvy, and the authors report that age, education, tech savviness, and privacy literacy all predicted participation. In other words, the recruitment funnel of a diary study is itself a sampling instrument, and it is a biased one.
The cadence decision matrix
Cadence is the single choice most likely to sink a diary study. Pick it from the research question, not the calendar.
| Research question | Cadence | Duration | Entry length | Main risk |
| How does a habit form after onboarding? | 1 entry/day, evening | 2-3 weeks | 2-3 minutes | Late-week decay |
| Where does a weekly workflow break? | Event-triggered, participant-initiated | 2 weeks | 3-5 minutes | Under-reporting of routine events |
| How does sentiment shift across a day? | 3-5 prompts spread across waking hours | 5-7 days | 60-90 seconds | Fatigue; skew toward available participants |
| What happens around one rare moment? | Event-triggered plus one weekly recap | 3-4 weeks | 5 minutes | Very sparse data |
| How does experience differ across markets? | 1 entry/day in the participant's own language | 2 weeks | 2-3 minutes | Translation drift in analysis |
Two rules follow from the evidence above. Spread prompts across the day rather than bunching them. And if you must choose between more prompts per day and more days, choose more days — decay inside a crowded day is steeper than decay across a well-paced week.
Where AI actually helps
The honest framing is that AI improves two specific seams and leaves the rest of the method unchanged.
Seam 1: the follow-up on each entry
The classic weakness of a diary entry is that it is thin. A participant writes "app was annoying today" and the researcher reads it four days later, too late to ask what happened. An AI moderator can ask one contextual follow-up in the moment — which is the difference between a log line and a usable observation.
Discipline matters here. One follow-up, not three. A diary entry that turns into an interview every evening is the fastest way to lose a panel, and NN/g's standing advice is to keep entries brief and stay in contact with participants to reduce fatigue and dropout.
Seam 2: synthesis across participants and days
A three-week study with 25 participants at one entry per day yields roughly 525 entries — small for statistics, awkward for manual coding, and exactly the size where analysis quietly gets skipped. AI-assisted thematic analysis suits this shape of corpus, provided a human owns the codebook and reviews the quotes. Our guide to analyzing open-ended responses applies the same pattern to survey text.
Where AI does not help
It does not fix a cadence that asks too much, it does not recover participants lost in onboarding, and it cannot tell you whether an absent entry means "nothing happened" or "I gave up." Missingness in diary data is informative, and interpreting it is a human job.
The Qualitati Diary Study Design Checklist
Work through this before recruiting. Each item maps to a failure that recurs in longitudinal qualitative designs.
- State the time-dependent question. If it could be answered in one interview, run the interview instead.
- Fix the cadence from the matrix above — and write down why, so it does not get renegotiated mid-fielding.
- Cap entry length at 3 minutes and pilot the timing with two people before launch.
- Avoid the 07:30-09:00 prompt window; anchor to midday or evening.
- Budget for onboarding loss, not just diary dropout. Assume the funnel from invitation to first entry is where most attrition happens.
- Pre-register the minimum viable dataset: the participants x days you need before analysis is meaningful.
- Write the entry prompt as a template, not as a fresh question each day, so entries stay comparable.
- Allow exactly one adaptive follow-up per entry.
- Plan the disclosure. If an AI moderates entries, participants should know before they consent.
- Decide the missingness rule in advance — what counts as a completed diary, and how partial diaries enter analysis.
Entry prompt template
A reusable daily template, kept deliberately short:
- What did you do with [product/task] today? (open text)
- How did that go? (1-5, plus one sentence)
- Was there any moment that stood out - good or bad? (open text, optional)
- Adaptive follow-up: one probe generated from the answer above, asked only when the entry contains an evaluative statement without a reason.
The last line is the whole AI contribution. Keep it that narrow.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX, and customer insights teams. For a diary study, the relevant pieces are:
- Conversational surveys with AI-driven follow-ups and branching logic — the natural container for a short daily entry plus one adaptive probe. See what conversational surveys are.
- ThemeLens thematic analysis, a map-reduce pipeline running across up to 100 transcripts at once, mapping codes to research questions and synthesizing themes with participant-anchored quotes — the shape of work a multi-week diary corpus creates.
- QDA Workspace for inductive and deductive coding and codebook generation, when you want to own the codebook rather than accept a generated one.
- Multilingual research in 10 languages (English, Chinese, French, Norwegian, Dutch, German, Spanish, Portuguese, Japanese, Arabic), so a cross-market diary study can run in each participant's own language. See running multilingual interviews.
- Transparent pricing: a free tier with 30 credits on signup, no credit card required, and published per-credit usage rates on the pricing page.
Qualitati does not ship a dedicated mobile diary app with push notifications. If a design depends on native mobile prompting at fixed times of day, plan that layer separately.
Limitations and trade-offs
Three caveats deserve to be stated plainly.
Self-report is still self-report. Logging closer to the moment reduces recall bias; it does not remove social desirability, and knowing an entry is due every evening can change the behavior being measured.
The compliance figures come from adjacent literatures. The 78% average and the 17.82% completion rate above come from experience sampling in social and clinical research, not from UX diary studies with commercial incentives. Treat them as calibration for the shape of attrition, not as forecasts for a specific panel.
Non-random dropout is the real threat to validity. If the people who keep logging are the ones who already like the product, three weeks of enthusiastic entries will read like evidence. Compare the profile of completers against the starting sample before presenting a single theme. This is a methodological claim worth reviewing with a researcher on your team rather than accepting from any tool, ours included.
Who this is for - and when not to use it
Use a diary study when the question involves change, frequency, or context you cannot observe; when the experience is distributed across a week rather than concentrated in a session; or when recall is known to be unreliable.
Do not use one when you need an answer this week, when the behavior happens once and is easy to remember, when a usability question would be better answered by watching someone attempt a task, or when you cannot commit a human to reading entries as they arrive.
FAQ
How long should a diary study run?
Two to four weeks is the common range. Shorter than a week rarely captures change; longer than four weeks compounds attrition without proportional gain unless the phenomenon is genuinely slow.
How many participants do I need?
Diary studies are usually run with 10-30 participants, because each contributes many entries. Plan for meaningful loss between recruitment and first entry, and recruit against a completed-diary target rather than an invitation count.
What compliance rate should I expect?
Published ESM work averages about 78% of prompts answered, declining across days (Rintala et al., 2019). Commercial UX studies with clear incentives may do better; unincentivized ones often do worse.
Can AI moderate diary entries?
Yes, for the narrow job of asking one contextual follow-up per entry. Disclose it during consent, keep the probe count fixed, and have a human read entries during fielding rather than only at the end.
How is a diary study different from a longitudinal survey?
A longitudinal survey asks the same closed questions at intervals to measure change on fixed dimensions. A diary study collects open, in-context accounts and lets the dimensions emerge from the entries.
How do I analyze diary data without drowning in it?
Code as entries arrive rather than in one pass at the end, keep the codebook human-owned, and use AI-assisted thematic analysis for cross-participant synthesis - then verify every theme against anchored participant quotes.
Bottom line
A diary study is worth running when the research question has a clock in it. The method's failure mode is not analysis - it is compliance, and compliance is decided before fielding starts by cadence, entry length, prompt timing, and the onboarding funnel. AI earns its place in two narrow slots: one adaptive follow-up per entry, and synthesis at the end. Design the rest yourself.
Start free with 30 credits - no credit card required - and run a conversational-survey diary entry, an AI-moderated interview, or a ThemeLens thematic analysis project. Or view transparent pricing first.
This article is an independent editorial summary. Methodology claims reflect publicly available information as of August 11, 2026. Design choices for your own study should be reviewed by a qualified researcher.