arXiv:2605.08186eess.AScs.AI2026-05

为自回归生成模型重新定义熵最小化,统一理论框架提升多场景性能

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

论文配图:Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models
图 1 · 摘自论文原文
  • 提出针对自回归模型的统一熵最小化公式,分解为逐标记策略梯度与熵损失
  • 在20多个音效、口音和多语言场景中持续提升Whisper语音识别性能
  • 适合研究测试时自适应或语音识别鲁棒性优化的学者与工程师

通过熵最小化实现的测试时自适应(TTA)在分类任务中表现优异,但其在生成式自回归模型中的应用仍缺乏理论支撑。现有方法依赖不同启发式策略,如使用伪标签的教师强制或基于策略梯度的强化学习,缺乏统一数学基础。本文针对自回归模型推导出严格的熵最小化公式,证明其精确目标可分解为逐标记策略梯度损失与逐标记熵损失,并将先前方法重新解释为该统一框架的部分实现。以Whisper语音识别模型为实验平台,在超过20个多样化场景(包括噪声环境、口音差异及多语言设置)中验证,所提方法显著且一致地提升性能。

原文摘要 · Abstract (English)

Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressive models remains theoretically fragmented. Existing approaches typically rely on distinct heuristics, such as teacher forcing with pseudo labels or policy-gradient-based reinforcement learning, without a unified mathematical foundation. In this work, we resolve this discrepancy by deriving a rigorous formulation of EM tailored to autoregressive models. We show that the exact objective naturally decomposes into a token-level policy gradient loss and a token-level entropy loss, and we reinterpret prior methods as partial realizations of this unified formulation. Using Whisper ASR as a testbed, we demonstrate that our approach consistently improves performance across more than 20 diverse domains, including acoustic noise, accents, and multilingual settings.

自回归模型测试时自适应语音识别熵最小化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。