arXiv:2507.16835eess.AScs.CL2025-07被引 9

测试30万场AI面试,找出语音对话系统最佳组合。

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

  • 用大模型当裁判,评估语音转文字+大模型+文字转语音的组合表现。
  • 谷歌语音识别+GPT-4.1+Cartesia语音合成效果最好,用户满意度高。
  • 技术指标好不等于用户体验好,人机交互有更复杂的影响因素。

基于语音的对话AI系统越来越多采用语音转文字(STT)、大语言模型(LLM)和文字转语音(TTS)串联架构。本文基于超过30万场由AI主持的求职面试数据,对多种STT x LLM x TTS组合进行了大规模实证比较。采用基于大模型作为评判者的自动化评估框架,评估对话质量、技术准确性及技能评估能力。分析五种生产级配置发现,使用谷歌STT、GPT-4.1与Cartesia TTS的组合在客观指标和用户满意度上均优于其他方案。令人意外的是,客观质量指标与用户满意度相关性较弱,表明语音交互中的用户体验受技术性能以外的因素影响。研究为多模态对话系统组件选型提供实践指导,并建立了一套可验证的人机交互评估方法。

原文摘要 · Abstract (English)

Voice-based conversational AI systems increasingly rely on cascaded architectures that combine speech-to-text (STT), large language models (LLMs), and text-to-speech (TTS) components. We present a large-scale empirical comparison of STT x LLM x TTS stacks using data sampled from over 300,000 AI-conducted job interviews. We used an LLM-as-a-Judge automated evaluation framework to assess conversational quality, technical accuracy, and skill assessment capabilities. Our analysis of five production configurations reveals that a stack combining Google's STT, GPT-4.1, and Cartesia's TTS outperforms alternatives in both objective quality metrics and user satisfaction scores. Surprisingly, we find that objective quality metrics correlate weakly with user satisfaction scores, suggesting that user experience in voice-based AI systems depends on factors beyond technical performance. Our findings provide practical guidance for selecting components in multimodal conversations and contribute a validated evaluation methodology for human-AI interactions.

语音生成大模型人机交互评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。