提出数学框架解析数据增强与一致性正则关系,提升模型泛化与稳定性。
A Mathematics Framework of Artificial Shifted Population Risk and Its Further Understanding Related to Consistency Regularization
- 构建数据增强的数学框架,揭示其本质为一致性正则项。
- 证明早期训练中存在负向风险缺口,影响收敛性能。
- 提出缓解方法,在多种场景下均优于现有方案。
数据增强是提升深度神经网络泛化能力与鲁棒性的关键技术。尽管常用于扩充样本量并作为一致性正则项,但其与正则化之间的内在联系仍缺乏系统研究。本文提出更全面的数据增强数学框架,证明偏移总体的期望风险等于原始风险加上一个可解释为一致性正则项的差距项。该框架进一步揭示该差距在训练初期具有负面影响。为此,我们提出有效缓解策略。通过在标准训练、分布外及不平衡分类等多种场景下的实验验证,结果表明所提方法在泛化能力和收敛稳定性上均超越对比方法。代码已开源:https://github.com/ydlsfhll/ASPR。
原文摘要 · Abstract (English)
Data augmentation is an important technique in training deep neural networks as it enhances their ability to generalize and remain robust. While data augmentation is commonly used to expand the sample size and act as a consistency regularization term, there is a lack of research on the relationship between them. To address this gap, this paper introduces a more comprehensive mathematical framework for data augmentation. Through this framework, we establish that the expected risk of the shifted population is the sum of the original population risk and a gap term, which can be interpreted as a consistency regularization term. The paper also provides a theoretical understanding of this gap, highlighting its negative effects on the early stages of training. We also propose a method to mitigate these effects. To validate our approach, we conducted experiments using same data augmentation techniques and computing resources under several scenarios, including standard training, out-of-distribution, and imbalanced classification. The results demonstrate that our methods surpass compared methods under all scenarios in terms of generalization ability and convergence stability. We provide our code implementation at the following link: https://github.com/ydlsfhll/ASPR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。