arXiv:2510.15406cs.CL2025-10被引 11

测试语音大模型对口吃等言语障碍的鲁棒性,发现性能大幅下降。

VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency

  • 构建多维度评估框架VocalBench-DF,系统测试语音模型对口吃的影响
  • 22个主流语音大模型在口吃语音下性能显著下降,平均下降超30%
  • 识别出音素级处理和长上下文建模是主要瓶颈,适合无障碍语音技术研究者

尽管语音大语言模型(Speech-LLMs)在诸多应用中表现优异,但其鲁棒性尚未得到充分检验,尤其在面对言语不流畅性时。现有评估多依赖理想化输入,忽视了帕金森病等疾病相关的常见言语障碍。本文探讨当前语音大模型在与有言语障碍用户交互时的表现。为此,我们提出VocalBench-DF框架,用于系统评估跨多维度的言语不流畅性。对22个主流语音大模型的评估显示,其性能出现显著退化,表明真实应用场景下的准备度有限。进一步分析指出,音素级处理和长上下文建模是导致失败的主要瓶颈。从组件和流程层面增强识别与推理能力可显著提升鲁棒性。这些发现凸显了改进不流畅性处理方法、构建真正包容性语音大模型的迫切需求。

原文摘要 · Abstract (English)

While Speech Large Language Models (Speech-LLMs) show strong performance in many applications, their robustness is critically under-tested, especially to speech disfluency. Existing evaluations often rely on idealized inputs, overlooking common disfluencies, particularly those associated with conditions like Parkinson's disease. This work investigates whether current Speech-LLMs can maintain performance when interacting with users who have speech impairments. To facilitate this inquiry, we introduce VocalBench-DF, a framework for the systematic evaluation of disfluency across a multi-dimensional taxonomy. Our evaluation of 22 mainstream Speech-LLMs reveals substantial performance degradation, indicating that their real-world readiness is limited. Further analysis identifies phoneme-level processing and long-context modeling as primary bottlenecks responsible for these failures. Strengthening recognition and reasoning capability from components and pipelines can substantially improve robustness. These findings highlight the urgent need for new methods to improve disfluency handling and build truly inclusive Speech-LLMs

语音模型鲁棒性口吃识别无障碍技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。