专为失语症语音设计的轻量级识别框架,部署在边缘设备上。
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
- 混合标准与失语语音训练,提升模型泛化能力。
- 使用GPT-4优化噪声失语转录,提升监督质量。
- 在失语语音上降低30%以上错误率,适合临床应用。
本文提出AS-ASR,一个基于Whisper-tiny的轻量级失语症专用语音识别框架,专为边缘设备上的低资源部署设计。该方法引入一种混合训练策略,系统性地以不同比例结合标准语音与失语语音,实现稳健泛化;同时采用GPT-4驱动的参考文本增强方法,对噪声失语转录进行修正,提升监督信号质量。我们在多种数据混合配置与评估设置下进行了大量实验。结果表明,微调后的模型显著优于零样本基线,在失语语音上将词错误率(WER)降低超过30%,同时保持对标准语音的识别性能。所提框架为真实世界中障碍性语音识别提供了可扩展、高效的解决方案。
原文摘要 · Abstract (English)
This paper proposes AS-ASR, a lightweight aphasia-specific speech recognition framework based on Whisper-tiny, tailored for low-resource deployment on edge devices. Our approach introduces a hybrid training strategy that systematically combines standard and aphasic speech at varying ratios, enabling robust generalization, and a GPT-4-based reference enhancement method that refines noisy aphasic transcripts, improving supervision quality. We conduct extensive experiments across multiple data mixing configurations and evaluation settings. Results show that our fine-tuned model significantly outperforms the zero-shot baseline, reducing WER on aphasic speech by over 30% while preserving performance on standard speech. The proposed framework offers a scalable, efficient solution for real-world disordered speech recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。