大模型能可靠纠正低资源语音识别错误,研究发现其纠错能力真实有效。
Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

- 用大模型进行生成式纠错,提升低资源弗里斯兰语识别效果。
- 最佳模型性能超越理想基准,且在私有数据集上仍有效。
- 揭示了模型纠错模式,适合语音识别与大模型融合研究者参考。
近年来自动语音识别(ASR)取得显著进展,但低资源语言的性能仍受限。大语言模型(LLMs)在生成式错误纠正(GER)方面展现出潜力,但其在低资源场景下的有效性尚未充分探索。此外,数据污染对报告改进效果的影响仍不明确。本研究针对低资源弗里斯兰语开展基于大模型的GER研究,除使用公开语料库外,还构建并采用包含非公开文本的弗里斯兰语离线数据集以控制数据污染。结果表明,GER在多数设置下均提升了ASR性能,最佳GPT-5结果优于理想基线词错误率(oracle WER)。在离线数据集上仍获得可比提升,表明改进反映真实的纠错能力。进一步的详细错误分析揭示了模型的纠错模式。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) has improved substantially in recent years, yet performance remains limited for low-resource languages. Large language models (LLMs) have shown promise for improving ASR through generative error correction (GER), but their effectiveness in low-resource settings remains underexplored. In addition, it remains unclear to what extent data contamination influences the reported improvements in LLM-based GER. This study investigates LLM-based GER for low-resource Frisian. In addition to a public corpus, we construct and use a Frisian offline dataset with non-public texts for evaluation to control for potential data contamination. Results show that GER improves ASR performance in most settings, with the best GPT-5.1 results surpassing oracle WERs. Comparable gains on the offline dataset indicate that improvements reflect true correction ability. We further provide a detailed error analysis revealing model correction patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。