解决大模型语音识别幻觉与定制难问题,提升实际应用鲁棒性。
Index-ASR Technical Report
- 融合噪声与上下文数据训练大模型,增强输入感知能力
- 在多个公开及内部测试集上表现优异,减少冗余输出
- 适合需要精准关键词识别的工业级语音应用
近年来,自动语音识别(ASR)取得显著进展,主要得益于大语言模型(LLM)驱动的新范式。尽管现有基于LLM的ASR系统在多种开源基准上表现强劲,但仍存在两大关键缺陷:一是易产生幻觉,生成过长且重复的内容,与声学输入不一致;二是缺乏灵活、细粒度的上下文定制支持。为此,我们提出Index-ASR,一种大规模基于LLM的ASR系统,旨在同时提升鲁棒性并支持可定制的热词识别。其核心思想是将大语言模型与富含背景噪声和上下文信息的大规模训练数据相结合。实验结果表明,Index-ASR在开源基准和内部测试集上均表现良好,凸显其在真实场景中的鲁棒性与实用性。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) has witnessed remarkable progress in recent years, largely driven by the emergence of LLM-based ASR paradigm. Despite their strong performance on a variety of open-source benchmarks, existing LLM-based ASR systems still suffer from two critical limitations. First, they are prone to hallucination errors, often generating excessively long and repetitive outputs that are not well grounded in the acoustic input. Second, they provide limited support for flexible and fine-grained contextual customization. To address these challenges, we propose Index-ASR, a large-scale LLM-based ASR system designed to simultaneously enhance robustness and support customizable hotword recognition. The core idea of Index-ASR lies in the integration of LLM and large-scale training data enriched with background noise and contextual information. Experimental results show that our Index-ASR achieves strong performance on both open-source benchmarks and in-house test sets, highlighting its robustness and practicality for real-world ASR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。