arXiv:2509.01814cs.CLcs.AI2025-09被引 2

AI语音面试官在问卷调查中表现优于传统系统,但情感识别仍有短板。

Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts

  • 用大语言模型实现实时语音访谈,支持追问与逻辑跳转。
  • 实时转录错误率较高,情绪识别能力有限,影响质性数据质量。
  • 适合量化研究,质性研究需谨慎评估使用场景。

基于Transformer的大型语言模型推动了“AI面试官”的发展,可实时开展语音问卷调查。本文综述现有证据,评估其在定量与定性研究中的适用性。从输入输出性能(如语音识别、回答记录、情绪处理)和语言推理能力(如追问、澄清、分支逻辑处理)两个维度进行分析。实地研究表明,当前AI面试官在定量与定性数据收集中已超越传统交互式语音应答(IVR)系统,但实时转录错误率、情绪识别能力不足以及后续追问质量不均,表明其在定性研究中的实用性与采纳需视具体情境而定。

原文摘要 · Abstract (English)

Transformer-based Large Language Models (LLMs) have paved the way for "AI interviewers" that can administer voice-based surveys with respondents in real-time. This position paper reviews emerging evidence to understand when such AI interviewing systems are fit for purpose for collecting data within quantitative and qualitative research contexts. We evaluate the capabilities of AI interviewers as well as current Interactive Voice Response (IVR) systems across two dimensions: input/output performance (i.e., speech recognition, answer recording, emotion handling) and verbal reasoning (i.e., ability to probe, clarify, and handle branching logic). Field studies suggest that AI interviewers already exceed IVR capabilities for both quantitative and qualitative data collection, but real-time transcription error rates, limited emotion detection abilities, and uneven follow-up quality indicate that the utility, use and adoption of current AI interviewer technology may be context-dependent for qualitative data collection efforts.

AI面试官语音访谈定性研究语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。