Method Fidelity in AI-Conducted Organisation-Design Interviews: Instrument Development and Evaluation
Hector Benitez Ventura and Naomi Stanford · Latent Variables
Organisation design depends on interview evidence, and the capacity to gather it has long been limited by what a practitioner can personally conduct. AI-conducted interviewing raises a prior question, which is whether such an instrument enacts a named method with measurable fidelity. This paper develops and evaluates an AI-conducted voice interviewer against a twelve-move workflow-based organisation-design method, scoring each transcript move by move with attempted and landed behaviour separated and each judgement linked to transcript evidence. Across five reader-facing development stages, mean fidelity increased from 6.9 to 9.4 moves of twelve, follow-up of substantive answers from 42.7% to 78.2%, and respondent words per interview from 839 to 2,473. One specified move was attempted in none of 175 baseline interviews, which shows that a method component can be present in a protocol and absent from behaviour. Two blinded human coders independently scored 1,104 move decisions, and against the adjudicated standard the automatic scorer reached macro-averaged precision of .953 and recall of .943, with an intraclass correlation of .91 for total fidelity. Because sequential development could not separate individual changes from fielding order, a subsequent factorial evaluation varied three instrument features independently as a supporting attribution check. Follow-up defaults and runtime duration control were each associated with higher fidelity, by 1.05 and 1.35 moves, while a mechanical one-question constraint reduced multi-question turns sixfold without improving fidelity. That agreement establishes the reliability of the fidelity measure, not the accuracy or representativeness of respondent accounts.