提出改进版Mixup,让合成数据更保真且不崩溃
A Generalized Theory of Mixup for Structure-Preserving Synthetic Data
- 设计可调权重的Mixup,更好保持原始数据统计结构
- 理论证明新方法能稳定保持方差与分布特性
- 适合关注数据合成质量与模型长期稳定的研究者
Mixup是一种广泛使用的数据增强技术,通过插值数据点提升模型泛化能力。尽管效果显著,但其生成合成数据的统计特性缺乏深入研究。本文揭示了Mixup可能扭曲方差等关键统计属性,导致数据合成中的意外后果。为此,我们提出一种具有广义灵活权重机制的新Mixup方法,理论上给出了保持(协)方差和分布特性的条件。数值实验表明,该方法不仅能有效保留原始数据的统计特征,还能在重复合成中维持模型性能,缓解以往研究中出现的模型坍塌问题。
原文摘要 · Abstract (English)
Mixup is a widely adopted data augmentation technique known for enhancing the generalization of machine learning models by interpolating between data points. Despite its success and popularity, limited attention has been given to understanding the statistical properties of the synthetic data it generates. In this paper, we delve into the theoretical underpinnings of mixup, specifically its effects on the statistical structure of synthesized data. We demonstrate that while mixup improves model performance, it can distort key statistical properties such as variance, potentially leading to unintended consequences in data synthesis. To address this, we propose a novel mixup method that incorporates a generalized and flexible weighting scheme, better preserving the original data's structure. Through theoretical developments, we provide conditions under which our proposed method maintains the (co)variance and distributional properties of the original dataset. Numerical experiments confirm that the new approach not only preserves the statistical characteristics of the original data but also sustains model performance across repeated synthesis, alleviating concerns of model collapse identified in previous research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。