评测主流语音模型在人机交互中的表现差异
Talking to Robots: A Practical Examination of Speech Foundation Models for HRI Applications
- 在8个真实数据集上测试4种先进语音模型
- 不同场景下性能差异大,错误率最高达40%
- 适合关注语音识别公平性与鲁棒性的研究者
现实世界中的自动语音识别系统需应对硬件限制或环境噪声导致的音频退化问题,同时适应多样用户群体。在人机交互(HRI)场景中,这些挑战叠加形成独特而严峻的识别环境。本文在8个公开数据集上评估了4种前沿语音识别系统,覆盖6类难度维度:领域特定、口音、噪声、年龄差异、语言障碍及即兴表达。分析显示,尽管标准基准分数相似,各模型在实际表现、幻觉倾向和固有偏见方面存在显著差异。这些局限对人机交互影响深远,识别错误可能干扰任务执行、削弱用户信任并危及安全。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems in real-world settings need to handle imperfect audio, often degraded by hardware limitations or environmental noise, while accommodating diverse user groups. In human-robot interaction (HRI), these challenges intersect to create a uniquely challenging recognition environment. We evaluate four state-of-the-art ASR systems on eight publicly available datasets that capture six dimensions of difficulty: domain-specific, accented, noisy, age-variant, impaired, and spontaneous speech. Our analysis demonstrates significant variations in performance, hallucination tendencies, and inherent biases, despite similar scores on standard benchmarks. These limitations have serious implications for HRI, where recognition errors can interfere with task performance, user trust, and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。