用大模型纠正儿童对话的语音识别错误,效果有好有坏。
Large Language Models based ASR Error Correction for Child Conversations
- 用大模型对零样本和微调后的语音识别结果进行纠错
- 在零样本和CTC模型输出上纠错效果显著,提升准确率
- 对自回归模型(如Whisper)的纠错能力有限,尤其依赖上下文时
自动语音识别(ASR)近年来取得显著进展,但准确转录儿童语音仍是重大挑战。大语言模型(LLMs)在改进语音识别结果方面展现出潜力,但在儿童对话场景中的应用仍不充分。本研究探索了使用LLMs纠正儿童对话语音识别错误的方法。通过在两个儿童对话语音数据集上对零样本和微调后的ASR输出进行实验,我们发现LLMs能有效改善零样本及基于CTC的微调ASR输出的错误,但在引入上下文信息或处理微调后的自回归型ASR(如Whisper)输出时,纠错效果仍受限。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In this study, we explore the use of LLMs in correcting ASR errors for conversational child speech. We demonstrate the promises and challenges of LLMs through experiments on two children's conversational speech datasets with both zero-shot and fine-tuned ASR outputs. We find that while LLMs are helpful in correcting zero-shot ASR outputs and fine-tuned CTC-based ASR outputs, it remains challenging for LLMs to improve ASR performance when incorporating contextual information or when using fine-tuned autoregressive ASR (e.g., Whisper) outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。