Hector Benitez Ventura, Noah Alexander, and Yashraj Patel · Latent Variables · June 2026
Abstract
Organizational change research often detects weak adoption long before it can identify the barrier that would make the next intervention actionable. We ask whether an AI-conducted voice interview preserves enough diagnostic signal for blinded post-hoc recovery of which ADKAR readiness barrier (awareness, desire, knowledge, ability, or reinforcement) is behind a participant's account. In a controlled benchmark where five of six matched workplace accounts each weakened one ADKAR element while the sixth supported all, blinded scorers who saw neither condition nor ground truth recovered the engineered barrier in 20 of 40 deficit cases: 50.0% accuracy (Wilson 95% CI 35.2–64.8), well above the 20% five-label chance baseline (p < .001). Participants read one of the six accounts of a workplace move to a tool called Flowboard and completed a voice interview about it. Three independent scoring passes labeled blinded transcript packets against a prespecified codebook with high inter-scorer reliability (nominal α = 0.76 for the six-way barrier label, above the prespecified 0.67 floor). Signal was strongest for knowledge and ability; reinforcement was hardest, and control accounts showed some residual overdiagnosis of knowledge gaps. These findings indicate that controlled, AI-conducted conversational data can carry recoverable readiness-barrier signal above chance. They do not yet establish field diagnosis: human-coder validation and prediction of real adoption outcomes remain necessary, and we frame these as the next steps for this research program.
Highlights
- Across 40 deficit cases, blinded consensus recovered the engineered ADKAR barrier in 20 of 40 (50.0% accuracy; Wilson 95% CI 35.2–64.8), well above the 20% five-label chance baseline (exact binomial p < .001).
- Three independent model-based scoring passes over the same blinded packets reached high agreement — nominal α = 0.76 for the six-way barrier label, above the prespecified 0.67 reliability floor.
- Knowledge and ability were recovered most reliably; reinforcement was the weakest target (2 of 8 recovered), suggesting interviews need stronger time-line probes about what happened after initial use.
- The full-support control served as a specificity check: 10 of 14 control rows were labeled 'none', with the four false positives concentrated in knowledge — a residual bias toward reading sparse procedural detail as a knowledge gap.
- The result is framed as a serious first benchmark, not evidence that conversational interviews can yet diagnose real organizations; human-coder validation and field-outcome prediction remain the next steps.
The full methodology, results, tables, and references are in the PDF. © Latent Variables.