针对游戏语音识别错误,提出结合检索增强的纠错框架。
Game-Oriented ASR Error Correction via RAG-Enhanced LLM
- 用检索增强生成和大模型生成游戏语料,增强数据多样性
- 在真实游戏场景中将字符错误率降6.22%,句错误率降29.71%
- 适合需要高精度语音交互的游戏开发与语音助手团队
随着多人在线游戏的发展,实时语音通信对团队协作至关重要。然而,通用语音识别系统在应对短语、快速发言、游戏术语和噪声等游戏特有挑战时表现不佳,导致频繁出错。为此,我们提出面向游戏的ASR错误纠正框架GO-AEC,融合大语言模型、检索增强生成(RAG)以及基于大模型与语音合成技术的数据增强策略。GO-AEC包含数据增强、基于N-best候选的纠错机制以及动态游戏知识库。实验表明,该框架在真实游戏场景中使字符错误率降低6.22%,句错误率降低29.71%,显著提升语音识别准确率。
原文摘要 · Abstract (English)
With the rise of multiplayer online games, real-time voice communication is essential for team coordination. However, general ASR systems struggle with gaming-specific challenges like short phrases, rapid speech, jargon, and noise, leading to frequent errors. To address this, we propose the GO-AEC framework, which integrates large language models, Retrieval-Augmented Generation (RAG), and a data augmentation strategy using LLMs and TTS. GO-AEC includes data augmentation, N-best hypothesis-based correction, and a dynamic game knowledge base. Experiments show GO-AEC reduces character error rate by 6.22% and sentence error rate by 29.71%, significantly improving ASR accuracy in gaming scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。