通过声调与韵律协同,提升粤语低资源语音识别准确率
CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
- 结合强制对齐提取声学特征,用LoRA微调Whisper增强声调区分
- 指令微调Qwen-Audio实现韵律感知纠错,字错误率显著降低
- 适合粤语、方言及低资源语音识别研究者参考
自动语音识别(ASR)对语言无障碍至关重要,但低资源粤语因标注数据有限、存在六个词汇声调、声调变调及口音差异而难以处理。现有模型如Whisper常出现高词错误率。大音频-语言模型(LALM)虽具更强上下文推理能力,但仍需显式声调与韵律声学提示。我们提出CantoASR,一种协作式ASR-LALM纠错框架,整合强制对齐提取声学特征、LoRA微调Whisper以增强声调判别力,以及指令微调的Qwen-Audio实现韵律感知纠错。在自发粤语数据上的评估显示,其字错误率(CER)显著优于Whisper-Large-V3。结果表明,将声学线索与LALM推理结合,为低资源声调及方言语音识别提供可扩展策略。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) is critical for language accessibility, yet low-resource Cantonese remains challenging due to limited annotated data, six lexical tones, tone sandhi, and accent variation. Existing ASR models, such as Whisper, often suffer from high word error rates. Large audio-language models (LALMs), in contrast, can leverage broader contextual reasoning but still require explicit tonal and prosodic acoustic cues. We introduce CantoASR, a collaborative ASR-LALM error correction framework that integrates forced alignment for acoustic feature extraction, a LoRA-finetuned Whisper for improved tone discrimination, and an instruction-tuned Qwen-Audio for prosody-aware correction. Evaluations on spontaneous Cantonese data show substantial CER gains over Whisper-Large-V3. These findings suggest that integrating acoustic cues with LALM reasoning provides a scalable strategy for low-resource tonal and dialectal ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。