用无标签数据提升模型泛化能力,提出新度量方法
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data

- 从参数空间信息几何出发,定义可无标签计算的局部不一致性
- 该度量与泛化差距显著相关,且在多种场景下提升模型表现
- 适合半监督、自监督学习,无需标签即可优化泛化性能
估计泛化差距并设计提升泛化的优化方法对深度学习的理论与应用均至关重要。利用无标签数据在此类任务中具有显著优势。本文提出一种新的泛化度量——局部不一致性,基于神经网络参数空间的信息几何视角。其关键特性在于无需显式标签即可计算。理论层面,我们建立了局部不一致性与费舍尔信息矩阵及损失海森矩阵的联系。实验表明,局部不一致性与泛化差距显著相关。基于此,我们提出不一致性感知最小化(IAM),将局部不一致性引入训练目标。在标准监督学习中,IAM性能媲美已有方法如锐度感知最小化(SAM)。此外,IAM在半监督和自监督学习中也表现优异,此时局部不一致性由无标签数据计算得出。
原文摘要 · Abstract (English)
Estimating the generalization gap and developing optimization methods that improve generalization are crucial for deep learning models, for both theoretical understanding and practical applications. Leveraging unlabeled data for these purposes offers significant advantages in real-world scenarios. This paper introduces a novel generalization measure, local inconsistency, derived from an information-geometric perspective on the parameter space of neural networks. A key feature of local inconsistency is that it can be computed without explicit labels. We establish theoretical underpinnings by connecting local inconsistency to the Fisher information matrix and the loss Hessian. Empirically, we demonstrate that local inconsistency correlates with the generalization gap. Based on these findings, we propose Inconsistency-Aware Minimization (IAM), which incorporates local inconsistency into the training objective. We demonstrate that in standard supervised learning settings, IAM enhances generalization, achieving performance comparable to that of existing methods such as Sharpness-Aware Minimization. Furthermore, IAM exhibits efficacy in semi- and self-supervised learning scenarios, where the local inconsistency is computed from unlabeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。