评测8款商用语音识别系统对失语症语音的性能,提供可复用的辅助语音接口选型基准。
Benchmarking Commercial Speech Recognition and Multimodal Large Language Models on Dysarthric Speech: Severity-Stratified Baselines and Architecture-Specific Prompting Effects
- 按失语严重程度分层评估,使用真实患者数据构建个性化基准。
- 严重失语时所有系统WER超51%,大模型无明显优势,最佳系统仅达1-2%轻度误差。
- 特定提示词可降低部分模型误差,但效果因模型架构而异,适配辅助技术部署者。
基于语音的人机交互已成为访问智能系统的主要方式,但失语症患者因识别准确率持续偏低而被系统性排除。尽管自动语音识别(ASR)在正常语音上可实现低于5%的词错误率(WER),但在失语症语音上性能急剧下降;而多模态大语言模型(MLLM)对此类语音的零样本表现尚不明确。本研究在TORGO失语语音语料库上评估了八款商用语音转写服务:四款传统ASR系统(AssemblyAI、Whisper large-v3、Deepgram Nova-3、Nova-3 Medical)和四款基于MLLM的系统(GPT-4o、GPT-4o Mini、Gemini 2.5 Pro、Gemini 2.5 Flash),采用词汇准确率、语义保留度与成本延迟进行衡量。识别性能随失语严重程度递减:轻度失语时,领先系统达到1-2%的低单数字WER;严重失语时,所有系统均超过51% WER,MLLM未展现优于传统ASR的优势。四条件提示词消融实验显示架构特异性效应:OpenAI模型中,逐字转录提示显著减少非目标语言漂移,使GPT-4o的严重层级WER从60.1%降至52.9%,GPT-4o Mini从66.0%降至约55%;而Gemini模型未见一致提升,甚至出现退化。语义指标与WER高度相关,在总体上呈冗余,但可识别出词汇错误高但语义意图部分保留的案例。这些按严重程度分层、针对个体说话者的基准,为辅助语音接口的实证技术选型提供了可复用参考。
原文摘要 · Abstract (English)
Voice-based human-machine interaction has become a primary means of accessing intelligent systems, yet individuals with dysarthria are systematically excluded by persistent gaps in recognition accuracy. Although automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades sharply for dysarthric speakers, while the zero-shot behaviour of multimodal large language models (MLLMs) on such speech remains unclear. We evaluate eight commercial speech-to-text services on the TORGO dysarthric speech corpus: four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash), using lexical accuracy, semantic preservation, and cost-latency measures. Recognition degraded consistently with severity. Mild dysarthria reached low single-digit WER, around 1-2% for the leading systems, whereas severe dysarthria exceeded 51% WER for every system, with no MLLM advantage over conventional ASR under default settings. A four-condition prompt ablation showed architecture-specific effects: for the OpenAI models, verbatim-transcription prompts reduced severe-tier WER mainly by suppressing non-target-language drift, lowering GPT-4o from 60.1% to 52.9% and GPT-4o Mini from 66.0% to about 55%; Gemini models showed no consistent benefit and sometimes degraded. Semantic metrics correlated strongly with WER and were largely redundant in aggregate, but identified cases where communicative intent was partly preserved despite poor lexical accuracy. These severity-stratified, per-speaker baselines provide a reusable reference for evidence-based technology selection in assistive voice interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。