arXiv:2511.03256cs.LGcs.CV2025-11NeurIPS被引 2

解耦熵最小化,提升模型在噪声环境下的鲁棒性

Decoupled Entropy Minimization

  • 将熵最小化拆分为聚类聚合与梯度抑制两部分,揭示其内在机制
  • 提出AdaDEM,解决高置信样本奖励衰减与类别分布偏差问题
  • 适用于噪声和动态环境中的弱监督学习,性能优于经典方法

熵最小化(EM)有助于降低类别重叠、弥合域间差距并限制不确定性,但其潜力受限。为研究其内部机制,我们将经典EM重新表述并解耦为两个相反作用的部分:聚类聚合驱动因子(CADF)奖励主导类别并促使输出分布尖锐化,而梯度抑制校准器(GMC)基于预测概率惩罚高置信类别。进一步揭示了经典EM因耦合形式导致的局限性:1)奖励坍缩阻碍高置信样本的学习贡献;2)易类偏差导致输出分布与标签分布不一致。为此,我们提出自适应解耦熵最小化(AdaDEM),对CADF带来的奖励进行归一化,并用边际熵校准器(MEC)替代GMC。AdaDEM超越了DEM*(经典EM的上限变体),在多种存在噪声和动态变化的不完美监督学习任务中表现更优。

原文摘要 · Abstract (English)

Entropy Minimization (EM) is beneficial to reducing class overlap, bridging domain gap, and restricting uncertainty for various tasks in machine learning, yet its potential is limited. To study the internal mechanism of EM, we reformulate and decouple the classical EM into two parts with opposite effects: cluster aggregation driving factor (CADF) rewards dominant classes and prompts a peaked output distribution, while gradient mitigation calibrator (GMC) penalizes high-confidence classes based on predicted probabilities. Furthermore, we reveal the limitations of classical EM caused by its coupled formulation: 1) reward collapse impedes the contribution of high-certainty samples in the learning process, and 2) easy-class bias induces misalignment between output distribution and label distribution. To address these issues, we propose Adaptive Decoupled Entropy Minimization (AdaDEM), which normalizes the reward brought from CADF and employs a marginal entropy calibrator (MEC) to replace GMC. AdaDEM outperforms DEM*, an upper-bound variant of classical EM, and achieves superior performance across various imperfectly supervised learning tasks in noisy and dynamic environments.

熵最小化弱监督鲁棒学习解耦机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。