Skip to content
Cati.botResearch at the speed of sound.
Cati.bot
Research notes

AI interviewing

Evaluating AI voice interviewers for survey research

An AI voice interviewer can conduct a fluent conversation and still produce a poor survey record. A qualified answer may be assigned to a category the respondent did not choose. A pause may be mistaken for the end of a turn. An explanation intended to help may change the meaning of the question. A demonstration lets a researcher hear how the system speaks; assessing these errors requires a study of how it interviews.

In conventional computer-assisted telephone interviewing, software manages the questionnaire while a person conducts the exchange. A voice AI system also takes responsibility for recognising speech, deciding how to respond and producing spoken output. Each responsibility needs an explicit specification. Allowing a model to choose a neutral clarification is a different intervention from allowing it to rewrite a question or infer an answer. Describing all three as automation conceals decisions that affect the resulting data.

What the published studies establish

A useful telephone pilot is the 2025 preprint by Leybzon and colleagues at VKL Research and SSRS. Two waves drew on the same pool of 104 SSRS Opinion Panel members: previous participants who had selected telephone interviewing and had not opted out of automated surveys. In the second wave, 70 people answered a call and 30 completed an interview. These were English interviews with an established, incentivised panel; the findings do not directly describe recruitment from a fresh telephone sample.

Completion improved between waves, but the software, questionnaire and calling schedule all changed. The second wave also included a shorter questionnaire. That design cannot isolate the effect of a software improvement. The subsequent human interviews asked about respondents’ experiences; they were not a control group answering the same substantive questionnaire. The study demonstrates that the system could administer interviews in this setting. It leaves comparative measurement quality unresolved.

A separate 2025 preprint by Lang and Eskenazi describes a deployment in Peru, with a sample consisting mainly of university students. The authors explicitly report that there was no human interviewer control group and that their analysis covered completed interviews, without feedback from people who dropped out. Those limitations matter when interpreting the paper’s favourable comparisons with human interviewing. Both telephone studies involved the systems’ developers; independent replications would strengthen the evidence.

Evidence from text interviewing addresses some related questions. In a peer-reviewed, 1,800-participant web experiment, published in August 2026, Barari and colleagues tested chatbots that probed answers and coded them during the interview. Open responses became more detailed and informative, with a small cost to respondent experience. Respondents’ tendency to agree with proposed interpretations also contributed to false positive codes. This is evidence about text interaction. Telephone interviewing adds speech recognition, interruptions and timing, which that experiment did not test.

Follow an answer through the system

Consider a hypothetical question about employment in the previous seven days. A respondent says, “I normally work, but I was on unpaid leave all last week.” The instrument’s definition determines the correct code. A system that hears only the opening clause may make a recognition or turn-taking error. A system that captures the whole sentence but applies the wrong reference period makes an interpretation error. If it then asks a question reserved for a different employment category, the error propagates through the interview.

An audit should therefore compare the available audio, transcript, assigned code and subsequent route. A tidy transcript does not establish that the code was justified. Nor does a valid code establish that the respondent understood the question. Include corrections, refusals, uncertain answers and requests for clarification in testing, alongside straightforward responses. State the expected action for each case before inspecting the system’s output.

Performance also needs to be examined across the voices a study expects to encounter. In a 2020 evaluation of five commercial speech recognition systems, Koenecke and colleagues found average word error rates of 35% for Black speakers and 19% for white speakers in their US speech samples. Those figures describe the systems and recordings tested then; they are not estimates for a current product. They give a concrete reason to examine subgroup performance using the deployed system and the intended population. For survey purposes, losing a negation or mishearing an eligibility answer can matter more than several errors in incidental words.

A comparison that can support a decision

Start by specifying the decision the pilot must inform. Replacing a human interviewer in an established tracker requires evidence about continuity of estimates. Adding an automated callback option requires evidence about additional participation and the characteristics of people it brings into the study. A technical rehearsal can identify broken routes in either case, but it cannot answer the population-level question.

For a replacement study, a defensible starting design is to randomly assign sampled cases to human or AI administration before contact. Keep the questionnaire, sample source, calling windows, attempt limits and incentives comparable, and document necessary differences in the introduction. The comparison then evaluates the administration protocols as respondents actually encounter them, including disclosure.

Preserve the assignment for analysis. If the AI introduction causes a particular group to refuse, comparing only completed interviews selects different people in the two arms. Differences between their answers may reflect participation, measurement, or both. Report outcomes for all assigned cases, then examine item responses and respondent composition separately. Where suitable external records exist, they can help assess accuracy; agreement with human CATI alone does not establish it.

Agree in advance which differences would change the commissioning decision. These might concern a primary prevalence estimate, break-offs at a sensitive section, or coding accuracy on a crucial eligibility item. A small pilot may have too little precision to distinguish an acceptable difference from a consequential one. As Lakens explains in his introduction to equivalence testing, a nonsignificant difference does not establish equivalence. Set substantively defensible bounds and plan sufficient precision to evaluate them.

Manual review needs its own sampling plan. Review a random selection of interactions as well as calls flagged for problems; flags alone cannot reveal failures the system has not learned to detect. Include partial interviews and failed exchanges. Have reviewers code against the instrument independently, reconcile disagreements, and report the denominator behind each error rate. Five routing failures among 100 opportunities convey more than five failures somewhere in a fieldwork archive.

Document the version that was evaluated

A finding applies to a particular configuration. A change in speech recognition, voice, instructions or conversational logic may change respondent behaviour even when the questionnaire file stays the same. Keep a version history linked to fieldwork dates, and rerun relevant test cases when components change. If a provider cannot identify an underlying model version, record that limitation.

AAPOR’s 2026 report on responsible AI integration recommends evaluating AI interviewing for validity, performance, sensitivity and reliability, alongside documenting AI’s role in the research. For a commissioning team, the practical deliverable is a methods record that connects the instrument, system configuration, fieldwork outcomes and validation results. Specify how respondents are informed, what material is retained for review, who can access it and when it is deleted. Recordings require a deliberate data-handling arrangement; their availability should never be assumed.

Finally, cost the protocol that passed evaluation. Include scripting, telephony, model usage, monitoring, manual adjudication, recontacts and unusable interviews. Report both total expenditure and cost per interview meeting prespecified quality criteria. Automation may lower some of these costs, but it does not by itself resolve nonresponse or the limits of statistical adjustment. Adoption is defensible when the study can meet its measurement and participation requirements at an acceptable cost. That conclusion has to be earned for the intended use.

Planning an AI interviewing pilot?

Tell the Voxworks team about your sample, questionnaire and validation requirements to discuss whether Cati.bot could fit the study.

Get in touch