arXiv:2508.08027cs.SDcs.AI2025-08被引 5

用大模型提升发音障碍者的语音识别准确率

Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches

  • 用自监督模型+大模型解码,修复发音扭曲
  • 大模型解码使识别错误率降低23.7%
  • 适合语音识别研究者和康复科技开发者

语音识别(ASR)在发音障碍者中因音素失真和高变异性而面临挑战。尽管自监督模型如Wav2Vec、HuBERT和Whisper表现出潜力,但其在发音障碍语音中的有效性尚不明确。本研究系统性地评估了这些模型与不同解码策略(包括CTC、seq2seq及基于大语言模型的解码,如BART、GPT-2、Vicuna)的组合效果。主要贡献包括:(1)构建发音障碍语音的ASR基准;(2)引入大语言模型解码以提升可理解性;(3)分析跨数据集泛化能力;(4)剖析不同严重程度下的识别错误模式。结果表明,大语言模型增强解码能通过语言约束实现音素恢复和语法修正,显著提升识别性能。

原文摘要 · Abstract (English)

Speech Recognition (ASR) due to phoneme distortions and high variability. While self-supervised ASR models like Wav2Vec, HuBERT, and Whisper have shown promise, their effectiveness in dysarthric speech remains unclear. This study systematically benchmarks these models with different decoding strategies, including CTC, seq2seq, and LLM-enhanced decoding (BART,GPT-2, Vicuna). Our contributions include (1) benchmarking ASR architectures for dysarthric speech, (2) introducing LLM-based decoding to improve intelligibility, (3) analyzing generalization across datasets, and (4) providing insights into recognition errors across severity levels. Findings highlight that LLM-enhanced decoding improves dysarthric ASR by leveraging linguistic constraints for phoneme restoration and grammatical correction.

语音识别大模型发音障碍自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。