arXiv:2505.00404cs.RO2025-05被引 1

提出通用中间层监督与正则化方法,提升模型泛化能力

iMacHSR: Intermediate Multi-Access Heterogeneous Supervision and Regularization Scheme Toward Architecture-Agnostic Training

  • 按通用标准选择中间层,用不同损失函数引导层次化表征学习
  • 在多个数据集上使mIoU提升最高达9.19%,超越单点输出监督
  • 适合需要跨架构训练的视觉理解任务,尤其关注泛化性能

深度监督通过在中间层施加辅助损失来增强训练效果,但存在三大未被充分探索的问题:(I) 现有方法高度依赖特定模型结构,缺乏通用性;(II) 中间层与输出层使用相同损失函数,导致中间层过早关注输出特异性特征,限制了可泛化的表示能力;(III) 隐层激活缺乏正则化,易产生过度自信预测,降低对未知场景的泛化性能。为此,本文提出一种面向架构无关的中间多接入异构监督与正则化方案(iMacHSR)。具体包括:(I) 基于预定义的架构无关标准选取多个中间层;(II) 在这些层上应用不同于输出层的损失函数,引导其学习多样且分层的表征;(III) 对所选层的隐层特征引入负熵正则化,抑制过度自信预测并缓解过拟合。这些中间项与输出层损失整合为统一优化目标,实现网络层级上的全面优化。以语义理解任务为例,在多个模型架构和数据集上验证iMacHSR,实验表明其相比传统单点输出监督方法,最高可提升mIoU达9.19%。

原文摘要 · Abstract (English)

While deep supervision is a powerful training strategy by supervising intermediate layers with auxiliary losses, it faces three underexplored problems: (I) Existing deep supervision techniques are generally bond with specific model architectures strictly, lacking generality. (II) The identical loss function for intermediate and output layers causes intermediate layers to prioritize output-specific features prematurely, limiting generalizable representations. (III) Lacking regularization on hidden activations risks overconfident predictions, reducing generalization to unseen scenarios. To tackle these challenges, we propose an architecture-agnostic, intermediate Multi-access Heterogeneous Supervision and Regularization (iMacHSR) scheme. Specifically, the proposed iMacHSR introduces below integral strategies: (I) we select multiple intermediate layers based on predefined architecture-agnostic standards; (II) loss functions (different from output-layer loss) are applied to those selected intermediate layers, which can guide intermediate layers to learn diverse and hierarchical representations; and (III) negative entropy regularization on selected layers' hidden features discourages overconfident predictions and mitigates overfitting. These intermediate terms are combined into the output-layer training loss to form a unified optimization objective, enabling comprehensive optimization across the network hierarchy. We then take the semantic understanding task as an example to assess iMacHSR and apply iMacHSR to several model architectures. Extensive experiments on multiple datasets demonstrate that iMacHSR outperforms conventional output-layer single-point supervision method up to 9.19% in mIoU.

深度监督泛化能力多层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。