arXiv:2608.11036cs.CL2026-08

构建缅甸语医学语音语料库并微调Whisper,提升临床对话识别准确率。

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

论文配图:myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR
图 1 · 摘自论文原文
  • 基于28小时本土语料,用FFT与LoRA微调Whisper模型。
  • 无增强时最佳模型WER达23.44%,优于更大通用模型。
  • 数据增强虽降清洁音质,但显著提升噪声/混响环境鲁棒性。

尽管Whisper模型得益于大规模多语言预训练,其在缅甸语医学语音上的表现仍有限。本文构建了一个由母语者录制并验证的高质量28小时缅甸语医学语音语料库,基于此对Whisper模型进行全量微调(FFT)和参数高效微调(PEFT,采用LoRA)。为评估鲁棒性,在波形和频谱层面施加受控噪声与模拟房间声学的数据增强。尽管增强会降低纯净语音性能,但在嘈杂和混响环境中显著提升了各微调方案的鲁棒性。最佳系统myMediWhisper-Medium(全量微调,无增强)在测试中达到23.44%的词错误率(WER),优于更大规模的通用领域微调模型。数据集及其他资源可在Huggingface仓库获取:https://huggingface.co/datasets/LULab/mediTalk-mm-rdy。

原文摘要 · Abstract (English)

Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers. We fine-tune Whisper models using full fine-tuning (FFT) and parameter-efficient fine-tuning (PEFT) with LoRA. To evaluate robustness, we apply waveform- and spectrogram-level data augmentation under controlled noise and simulated room acoustics. While augmentation reduces performance on clean speech, it significantly improves robustness in noisy and reverberant environments across FFT and PEFT settings. Our best-performing system, fully fine-tuned myMediWhisper-Medium without augmentation, achieves a state-of-the-art Word Error Rate (WER) of 23.44%, outperforming much larger general-domain fine-tuned models. Dataset and other resources can be found at the Huggingface repository: https://huggingface.co/datasets/LULab/mediTalk-mm-rdy.

语音识别医学ASR缅甸语Whisper微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。