从信息论角度解析数据增强如何影响模型泛化与不变性学习。
The geometry of invariant learning: an information-theoretic analysis of data augmentation and generalization
- 构建信息论框架,将数据增强视为原数据分布与变换分布的复合。
- 推导出泛化误差新上界,包含分布差异、算法稳定性与增强敏感性三项。
- 引入群直径概念,揭示增强范围与泛化性能间的内在权衡。
数据增强是现代机器学习中提升泛化能力最常用的技术之一,常被解释为促进对标签无关变换的不变性。然而其理论作用仍不完全明晰。本文提出一种基于信息论的系统性框架,分析增强对泛化和不变性学习的影响。该框架建立在互信息相关的泛化界基础上,将增强后的数据分布建模为原始数据分布与变换分布的复合,自然导出轨道平均损失函数。在损失函数与增强过程满足弱子高斯假设的前提下,我们推导出新的泛化界,将其分解为三项可解释成分:(1)原始数据与增强数据间的分布差异;(2)算法对训练数据的依赖程度(稳定性项);(3)增强可变性带来的敏感度。为进一步连接边界与增强群的几何结构,我们引入群直径概念——即增强在输入空间中所能诱导的最大扰动。群直径作为统一控制参数,同时约束三项误差,并揭示内在权衡:小直径保持数据保真但正则化不足,大直径提升稳定性却增加偏差与敏感度。数值实验验证了理论边界能可靠跟踪并预测真实泛化差距。
原文摘要 · Abstract (English)
Data augmentation is one of the most widely used techniques to improve generalization in modern machine learning, often justified by its ability to promote invariance to label-irrelevant transformations. However, its theoretical role remains only partially understood. In this work, we propose an information-theoretic framework that systematically accounts for the effect of augmentation on generalization and invariance learning. Our approach builds upon mutual information-based bounds, which relate the generalization gap to the amount of information a learning algorithm retains about its training data. We extend this framework by modeling the augmented distribution as a composition of the original data distribution with a distribution over transformations, which naturally induces an orbit-averaged loss function. Under mild sub-Gaussian assumptions on the loss function and the augmentation process, we derive a new generalization bound that decompose the expected generalization gap into three interpretable terms: (1) a distributional divergence between the original and augmented data, (2) a stability term measuring the algorithm dependence on training data, and (3) a sensitivity term capturing the effect of augmentation variability. To connect our bounds to the geometry of the augmentation group, we introduce the notion of group diameter, defined as the maximal perturbation that augmentations can induce in the input space. The group diameter provides a unified control parameter that bounds all three terms and highlights an intrinsic trade-off: small diameters preserve data fidelity but offer limited regularization, while large diameters enhance stability at the cost of increased bias and sensitivity. We validate our theoretical bounds with numerical experiments, demonstrating that it reliably tracks and predicts the behavior of the true generalization gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。