arXiv:2605.02782cs.AIcs.CL2026-05被引 2

测试语音模型能否利用临床信息提升口吃语音识别,发现现有模型基本无效。

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

  • 构建SAP数据集上的评测基准,测试诊断标签与临床描述对识别的影响。
  • 多数模型在加入临床信息后词错误率不降反升,改善微乎其微。
  • 用LoRA微调可实现52%相对降错,适合唐氏综合征等特定人群使用。

自动语音识别系统在口吃及其他异常语音上仍表现脆弱。近期的音频-语言模型有望通过推理时引入临床上下文来提升性能,但尚不清楚这些模型是否能有效利用此类信息。我们基于语音可访问性项目(Speech Accessibility Project, SAP)数据集构建了一个评测基准,检验诊断标签、临床医生评分及逐步丰富的临床描述能否提升口吃语音的转录准确率。在九个模型的匹配对比中,我们发现当前模型并未有效利用这些上下文:基于诊断的提示和临床详述提示带来的改进微乎其微,甚至常导致词错误率(WER)上升。我们进一步通过上下文依赖的微调实验表明,采用多种临床提示格式的LoRA适配可实现0.066的WER,相比冻结基线降低52%相对错误率,且在无上下文时保持性能。子组分析显示,唐氏综合征及轻度严重程度说话者获益显著。该结果揭示了当前模型的局限,并为更包容的语音识别提供了可测量的评估平台。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by conditioning on additional clinical context at inference time, but it is unclear whether these models can make use of such information. We introduce a benchmark built on the Speech Accessibility Project (SAP) dataset that tests whether diagnosis labels, clinician-derived speech ratings, and progressively richer clinical descriptions improve transcription accuracy for dysarthric speech. Across matched comparisons on nine models, we find that current models do not meaningfully use this context: diagnosis-informed and clinically detailed prompts yield negligible improvements and often degrade word error rate. We complement the prompting analysis with context-dependent fine-tuning, showing that LoRA adaptation with a mixture of clinical prompt formats achieves a WER of 0.066, a 52% relative reduction over the frozen baseline, while preserving performance when context is unavailable. Subgroup analyses reveal significant gains for Down syndrome and mild-severity speakers. These results clarify where current models fall short and provide a testbed for measuring progress toward more inclusive ASR.

语音识别口吃语音多模态上下文医疗辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。