How to Run Concept Testing Interviews (2026)
Qualitati Research Team · 2026-08-04 · 9 min read
Short answer: Concept testing interviews are moderated conversations in which participants react to a described or prototyped product concept and are probed on why. Run them monadically, recruit to the buying decision rather than the demographic, separate reaction from purchase intent, and treat stated preference as a hypothesis — not a forecast.
What are concept testing interviews?
Concept testing is the practice of putting an idea — a product, feature, positioning line, package, or ad — in front of the target audience before you build it, and observing how they engage with it. Contentsquare's concept testing guide frames it as gathering feedback on a new idea to judge viability and improve the final product.
Concept testing interviews are the qualitative half of that. A survey tells you that 62% of people rated a concept 4 or 5. An interview tells you that they rated it 4 because they misread the pricing model, which is a completely different product decision. Most practitioner guidance — including Qualtrics and Looppanel — recommends blending the two rather than choosing.
Who this is for
- Product managers deciding whether a roadmap bet survives contact with users.
- UX researchers validating a concept before design invests in high-fidelity work.
- Insights and brand teams screening messaging, packaging, or positioning routes.
- Founders pressure-testing a value proposition before writing code.
Key takeaways
- Concept testing interviews answer why a concept lands or fails; concept surveys answer how much. Run both, in that order.
- Use a monadic design by default — one concept per participant — unless your decision is explicitly a choice between routes.
- Recruit against the decision, not the persona. The person who would actually buy or approve is the only valid respondent.
- Stated purchase intent is systematically inflated. Treat it as a relative ranking signal, never an absolute forecast.
- Five to eight interviews per concept per segment is a reasonable exploratory floor; scale up when segments are heterogeneous.
Step 1: Write the decision before you write the guide
The most common failure in concept testing is running a study that cannot change anything. Before drafting a single question, write one sentence in this form:
"If participants respond [X], we will [do A]; if they respond [Y], we will [do B]."
If you cannot fill in both branches with actions your team would genuinely take, you are running a reassurance exercise, not research. Kill the study or reframe the decision.
Step 2: Choose a concept-test design
The design determines what your data can and cannot support. As of August 4, 2026, the three designs in common practitioner use trade off differently:
| Design | How it works | Best for | Main risk |
| Monadic | Each participant sees exactly one concept | Clean, uncontaminated read on a single concept | Needs more participants to compare routes |
| Sequential monadic | Each participant sees several concepts in rotated order | Comparing 2–4 routes with a smaller sample | Order and fatigue effects; later concepts get shallower reactions |
| Comparative / choice | Concepts are shown side by side | Explicit trade-off decisions and prioritization | Forces a preference that may not exist in the real market |
Default to monadic for qualitative work. Side-by-side comparison manufactures differentiation: participants will find a distinction between two concepts even when a real customer, encountering either one alone in the market, would not notice one. Decision Analyst's long-standing concept testing whitepaper makes a related argument about the "uniqueness" trap — novelty in a test environment is not the same as differentiation in a purchase environment.
Step 3: Recruit against the decision, not the demographic
"UX researchers at B2B SaaS companies, 25–45" is a demographic. "People who have chosen a research tool for their team in the last 12 months" is a decision. The second screener yields concept feedback you can act on; the first yields opinions.
Practical screener rules:
- Screen on recent behavior, not on claimed interest. "Have you done X in the last 90 days" beats "would you be interested in X."
- Include at least one segment that should reject the concept. A concept everyone likes has usually been described too vaguely to disagree with.
- Exclude professional respondents and anyone who can name your company as the sponsor.
- Ensure the respondent pool actually reflects the target audience rather than whoever was convenient — unrepresentative feedback is the standard route to a costly wrong conclusion.
Step 4: Present the concept without selling it
Concept stimulus should be neutral, complete, and boring. Three rules:
- No adjectives you would not put in a legal document. "Seamless," "effortless," and "powerful" contaminate the reaction. Describe what it does.
- State the price, or state that price is not yet set. A concept without a price is tested against an imagined price of zero, and reactions to it are worthless for a build decision.
- Hold fidelity constant across concepts. A polished mockup will beat a text description regardless of the underlying idea.
Step 5: Run the probe ladder
Here is the Qualitati Concept Probe Ladder — a five-rung sequence for each concept. Each rung must be exhausted before the next; skipping to rung 4 is how teams end up with polite, useless data.
| Rung | Goal | Example prompt |
| 1. Comprehension | Confirm they understood the concept, unaided | "In your own words, what is this and who is it for?" |
| 2. Relevance | Locate the concept in their actual life or workflow | "When was the last time you had the problem this addresses? Walk me through it." |
| 3. Reaction and reasoning | Get the evaluation and its basis | "What's your first reaction? What specifically is driving that?" |
| 4. Substitution | Identify the real competitive set | "If this existed tomorrow, what would you stop doing?" |
| 5. Cost of adoption | Surface the friction that kills good concepts | "Who else would have to agree before you could use this?" |
Rung 4 is the most underused and the most diagnostic. A concept with an enthusiastic reaction and no answer to "what would you stop doing" is a concept with no budget behind it.
Questions to avoid
- "Would you use this?" — predicts almost nothing; ask what they do today instead.
- "How much would you pay?" as an open question — yields anchoring noise. Test specific prices.
- "Do you like the blue or the green?" — preference questions on details before the concept itself has been validated.
- Any question that contains its own answer ("How helpful would it be to have...").
Step 6: Analyze reactions, not ratings
The analytic unit in a concept test interview is the reason, not the score. Code every transcript for four things:
- Misreads — what participants thought the concept was, when they were wrong. Cluster these; they are usually a copy problem, not a product problem.
- Objections — distinguish objections to the concept from objections to the description. Only the first should change the roadmap.
- Substitution targets — what the concept would displace, in the participant's words.
- Adoption blockers — approvals, integrations, switching costs, contracts.
Then run the counts. If seven of eight participants misread the same sentence, that is a finding with far more decision value than a mean rating of 3.8.
The Qualitati Concept Test Readiness Scorecard
Score your study out of 10 before fielding. Anything below 7 will produce data your team argues about rather than acts on.
| # | Criterion | Pass condition |
| 1 | Decision framing | Both branches of the if/then decision are written and both are acceptable to the team |
| 2 | Design fit | Monadic unless the decision is genuinely a choice between routes |
| 3 | Screener validity | Screens on behavior in a defined recent window, not on stated interest |
| 4 | Disconfirming segment | At least one segment included that could plausibly reject the concept |
| 5 | Stimulus neutrality | No promotional adjectives; fidelity held constant across concepts |
| 6 | Price present | A price or explicit price placeholder is shown |
| 7 | Probe depth | Guide reaches rung 4 (substitution) for every concept |
| 8 | Sample floor | ≥ 5–8 participants per concept per distinct segment |
| 9 | Analysis plan | Coding frame written before fielding, not after |
| 10 | Quant follow-up | A plan exists to size any qualitative signal that would move a large investment |
How many concept testing interviews do you need?
For exploratory concept work, five to eight interviews per concept per meaningfully distinct segment is a defensible starting point. The logic borrows from usability research, where Nielsen Norman Group's five-user guidance rests on diminishing returns within a homogeneous user group — and NN/g is explicit that the rule is a starting point for iteration and does not apply to quantitative measurement or to heterogeneous audiences.
That caveat matters more in concept testing than in usability testing. Usability problems are properties of the interface and repeat across users. Concept reactions are properties of the person's situation and do not. If you have three buyer types, you have three studies, not one.
Where Qualitati fits
Qualitati is an AI user research platform for product, UX, and customer insights teams. For concept testing specifically:
- AI-moderated interviews in text or voice run the probe ladder consistently across every participant, so rung 4 does not get skipped when a session runs long.
- Conversational surveys with AI-driven follow-ups and branching logic let you field a concept test at survey scale while still capturing the why behind each rating.
- Active Listener mode supports a human moderator with real-time prompts and section tracking when the concept is sensitive enough to warrant a person in the room.
- ThemeLens runs thematic analysis across up to 100 transcripts at once, mapping codes to research questions — which is how you get misread and objection clusters out of a monadic study quickly.
- QDA Workspace supports deductive coding against the four-code frame above, plus inductive coding for what you did not anticipate.
- Multilingual research in 10 languages, for concepts being tested across markets.
Pricing is published per credit, and new accounts start with 30 credits, no credit card required. View transparent pricing or start free with 30 credits.
Limitations and when not to use this approach
Concept testing interviews have real methodological ceilings, and it is worth being blunt about them.
- Stated preference is not revealed preference. People are poor forecasters of their own future behavior, and interview settings amplify agreeableness. Concept interviews rank and explain; they do not predict adoption rates.
- Novelty inflates scores. Anything new tests well relative to the familiar. Compare against a benchmark concept you have previously fielded, not against zero.
- Small qualitative samples cannot size a market. If the decision involves significant investment, follow the interviews with a properly powered quantitative test.
- Don't use concept interviews for pricing points. Use a dedicated pricing method; open-ended willingness-to-pay questions in interviews produce anchoring artifacts.
- Don't use them when the concept is unbuildable. Testing something engineering has already ruled out wastes participant goodwill and team credibility.
- AI moderation caveat: an AI moderator holds the guide consistently but does not read a room. For emotionally loaded or highly sensitive concepts, a human moderator — optionally supported by Active Listener — remains the better instrument. Any methodology claim that would carry weight in a published study should be reviewed by a human researcher before use.
Bottom line
Concept testing interviews earn their cost when they are built around a decision, run monadically, recruited on behavior, and probed all the way to substitution and adoption cost. Run them before the quantitative test, code the reasons rather than the ratings, and hold your stated-intent numbers loosely.
FAQ
What is the difference between concept testing and usability testing?
Concept testing evaluates whether an idea is worth building; usability testing evaluates whether a built thing can be used. Concept testing happens before design investment, usability testing after. The methods, sample logic, and success criteria all differ.
How many concepts can I test in one interview?
Two to four in a sequential monadic design, with the order rotated across participants. Beyond four, fatigue flattens the later reactions and the data stops being comparable.
Can AI moderate concept testing interviews?
Yes. AI-moderated interviews run a consistent probe sequence with adaptive follow-ups across many participants at once, which suits concept testing well because the guide is highly structured. Sensitive or emotionally complex concepts still benefit from a human moderator.
Should I show a price during concept testing?
Yes, or explicitly state that pricing is not yet set. A concept evaluated without a price is implicitly evaluated at zero cost, which inflates every downstream signal.
Is qualitative concept testing enough on its own?
For direction and diagnosis, often yes. For go/no-go decisions carrying significant investment, no — follow up with a quantitative concept test sized to detect the difference that would change your decision.
What sample size do I need per concept?
Five to eight participants per concept per distinct segment is a reasonable exploratory floor. Segment heterogeneity, not total headcount, is what drives the real requirement.
Get started
Run an AI-moderated concept test, a conversational survey, or a thematic analysis project on Qualitati. Start free with 30 credits — no credit card required — or compare transparent per-credit pricing. See also our guides to writing an AI-moderated discussion guide, building a screener for AI-moderated interviews, and running continuous discovery interviews.
Last updated: August 4, 2026. This guide is an independent editorial summary of publicly available methodological guidance; product claims describe Qualitati only.