| AI Interviewer | Leading questions | 0% Across all scored turns in a July 2026 internal evaluation with English and Chinese participant personas, the AI interviewer asked no leading questions — questions that presuppose or suggest a desired answer. Method: Simulated interviews against the production prompt pipeline; each turn scored by an LLM judge with rule-based checks, spot-checked by a researcher. |
|---|
| AI Interviewer | Probing on shallow answers | 83% When a participant gave a shallow or one-line answer, the interviewer followed up with a probe (rather than moving on) in 83% of opportunities. Method: Same July 2026 evaluation; probe opportunities identified from turn context, follow-up behaviour scored per opportunity. |
|---|
| ThemeLens (thematic analysis) | Quote verification | 100% of displayed quotes verified Every supporting quote shown in a ThemeLens report is checked against the source transcript before display. Quotes that cannot be located are dropped; quotes attributed to the wrong participant are reassigned. Nothing unverifiable is shown. Method: A verification pass runs inside the analysis pipeline itself (normalized text matching, ellipsis-aware). Kept / reassigned / dropped counts are recorded for every analysis. |
|---|
| QDA Workspace (AI coding) | Segment-offset fidelity | 100% On our ground-truth coding benchmark, every AI-coded segment now maps to the exact character span in the source document — up from 0% before the July 2026 offset-resolution work. Codes point at the text they actually describe. Method: Fixture transcripts with known correct segments; automated comparison of returned offsets against ground truth. |
|---|
| QDA Workspace (reliability) | Inter-rater reliability | Cohen's kappa, code-level When two independent AI coders (different model families) code the same document, agreement is reported with code-level Cohen's kappa — the same statistic reviewers expect from human double-coding — instead of a vague match percentage. Method: Segment pairs matched by span overlap; kappa computed on code assignment; coverage-based kappa reported when segmentation differs. |
|---|
| Annotator (batch coding) | Inter-coder reliability on every job | Krippendorff's α, Fleiss' κ, Cohen's κ Every multi-model annotation job returns reliability statistics across the models that coded it, with per-column agreement heatmaps, so you can see exactly where models disagree before trusting a column. Method: Standard reliability estimators computed over the models' parallel codings of identical rows. |
|---|
| Voice Analytics | Honest confidence labelling | Per-speaker calibration Emotion labels are computed relative to each speaker's own vocal baseline rather than absolute thresholds, and labels that cannot be calibrated (too few turns, flat delivery) are explicitly marked low-confidence instead of over-claimed. Method: Robust per-speaker baselines (median/IQR z-scores); emotion-bearing words in the transcript anchor vocal expressiveness; text–voice incongruence is flagged for human re-listening. |
|---|
| Digital Twin Panel | Grounded answers + enforced consent | Verbatim grounding Twin answers are grounded in the participant's actual interview turns retrieved for each question, calibration honestly reports when source data is insufficient, and twins cannot be activated or queried without explicit participant consent — enforced server-side and audited. Method: Retrieval over source transcripts (English and Chinese); calibration status computed from grounding coverage; consent checks at activation and at every simulated interaction. |
|---|