构建真实语音诊断基准,揭示语音识别在复杂场景下的失效机制。
Back to Basics: Revisiting ASR in the Age of Voice Agents
- 基于真实人类语音构建多语言诊断基准,分解环境、人群与语言三类退化因素。
- 七款主流模型在真实场景中表现严重不均,部分条件下错误率超60%。
- 发现模型会生成看似合理却未说出的内容,带来实际应用安全风险。
自动语音识别(ASR)系统在精心筛选的基准测试中已接近人类水平,但在真实语音代理场景中仍频繁失效,而现有评估未能系统覆盖这些条件。由于缺乏能隔离具体失败因素的诊断工具,从业者无法预判哪些语言、在何种条件下会导致性能下降。本文提出WildASR,一个源自真实人类语音的多语言(四语种)诊断基准,从环境退化、人口差异和语言多样性三个维度分解ASR鲁棒性。评估七种广泛使用的ASR系统后发现,性能严重且不均衡下降,模型鲁棒性无法跨语言或条件迁移。关键的是,当输入部分或受损时,模型常生成看似合理但从未被说出的内容,对下游代理行为构成实际安全风险。结果表明,针对特定因素的隔离式评估对理解并提升生产环境中ASR可靠性至关重要。除基准外,我们还提供三个可供从业者用于部署决策的分析工具。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without diagnostic tools that isolate specific failure factors, practitioners cannot anticipate which conditions, in which languages, will cause what degree of degradation. We introduce WildASR, a multilingual (four-language) diagnostic benchmark sourced entirely from real human speech that factorizes ASR robustness along three axes: environmental degradation, demographic shift, and linguistic diversity. Evaluating seven widely used ASR systems, we find severe and uneven performance degradation, and model robustness does not transfer across languages or conditions. Critically, models often hallucinate plausible but unspoken content under partial or degraded inputs, creating concrete safety risks for downstream agent behavior. Our results demonstrate that targeted, factor-isolated evaluation is essential for understanding and improving ASR reliability in production systems. Besides the benchmark itself, we also present three analytical tools that practitioners can use to guide deployment decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。