arXiv:2509.12508cs.CLcs.AI2025-09被引 20

Fun-ASR融合大模型与强化学习,提升真实场景下语音识别的准确率与稳定性。

Fun-ASR Technical Report

  • 结合海量数据、大模型和强化学习,优化语音识别流程。
  • 在真实工业评估集上表现超越现有系统,显著提升实用性能。
  • 针对流式传输、噪声鲁棒性等实际需求深度优化,适合落地应用。

近年来,自动语音识别(ASR)在数据规模扩大、模型容量增长以及与大语言模型(LLM)深度融合的推动下取得了突破性进展。然而,LLM容易产生幻觉,严重影响真实应用场景中的用户体验。本文提出 Fun-ASR,一个基于大规模数据、大模型容量、LLM集成与强化学习协同优化的 ASR 系统,在多样复杂的语音识别场景中达到业界领先性能。该系统特别针对实际部署优化,涵盖流式处理能力、噪声鲁棒性、多语言混用支持、热词定制等功能,满足真实业务需求。实验表明,尽管多数基于 LLM 的 ASR 系统在开源基准上表现优异,但在真实工业评估集上往往表现不佳;而 Fun-ASR 在真实应用数据集上实现最优表现,验证了其在实际环境中的有效性与鲁棒性。代码与模型已开源:https://github.com/FunAudioLLM/Fun-ASR。

原文摘要 · Abstract (English)

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR .

语音识别大模型工业落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。