arXiv:2604.08003eess.AScs.CL2026-04被引 1

通过熵分配视角优化语音大模型,提升识别精度并减少幻觉。

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

  • 从熵分配角度重新设计训练策略,增强语音与文本模态对齐。
  • 仅用23亿参数达到顶尖性能,显著降低幻觉率。
  • 适合关注低延迟、高鲁棒性的语音识别落地应用者。

将大语言模型(LLM)融入自动语音识别(ASR)已成为主流范式。尽管现有基于LLM的ASR模型在公开基准上表现优异,但在识别质量、延迟和计算开销之间仍难平衡,幻觉问题进一步制约实际部署。本文从熵分配视角重新审视基于LLM的ASR,提出三个指标以刻画不同训练范式在语音编码器与LLM间熵压缩的分配方式。为解决现有方法的熵分配效率低下问题,我们提出一种基于能力边界感知的分阶段训练策略,兼顾参数效率与幻觉鲁棒性。具体而言,重构预训练策略以缓解语音-文本模态差距,并引入对齐与联合微调之间的迭代异步SFT阶段,以保持功能解耦并约束编码器表示漂移。在中文和英文基准上的实验表明,本方法仅使用2.3B参数即达到与最先进模型相当的性能,同时通过解耦设计有效抑制幻觉。

原文摘要 · Abstract (English)

Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three metrics to characterize how training paradigms allocate entropy reduction between the speech encoder and the LLM. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a principled multi-stage training strategy grounded in capability-boundary awareness, optimizing parameter efficiency and hallucination robustness. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on Mandarin and English benchmarks show that our method achieves competitive performance with state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.

语音识别大模型熵分配幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。