提升数据选择效果:解决元训练中梯度噪声与特征不足问题
On the Difficulty of Learning a Meta-network for Training Data Selection
- 通过增大批次规模改善梯度信噪比,稳定元学习训练
- 在4个基准上平均提升5.49%,优于无选择和强基线
- 设计分布位置与训练动态特征,有效捕捉数据质量差异
合成数据被广泛用于训练神经网络,但与真实数据的分布差异限制了其效果。常用方法是通过双层优化学习数据权重,称为元学习训练数据选择(MTS)。然而实践中MTS常表现不佳。我们发现两个关键障碍:梯度信噪比(GSNR)低导致优化困难,以及缺乏与数据质量相关的有效特征。本文对MTS进行数学分析,揭示归一化数据权重的动态特性及数据质量差异与低GSNR的关系。分析表明,增大批次规模是简单有效的解决方案。此外,提出一组能捕捉训练数据分布位置和训练动态的有用特征。在四个基准上的实验显示一致改进,平均性能相比无选择提升5.49%,相比最强基线提升2.89%。
原文摘要 · Abstract (English)
Synthetic data are increasingly used to train neural networks, yet distributional mismatch with real data limits their effectiveness when used indiscriminately. A common strategy is to learn data weights via bi-level optimization, which we refer to as Meta-learning for Training-data Selection (MTS). Interestingly, in practice, MTS often performs below expectation. We identify two obstacles in properly training MTS: a poor gradient signal-to-noise ratio (GSNR), which causes optimization difficulties, and lack of informative features that correlates with data quality. We present a mathematical analysis of MTS, which reveals the dynamics of normalized data weights and the relation between disparate data quality and poor GSNR. The analysis suggests a a simple yet effective solution: increasing the batch size. Further, we propose a set of informative features that capture the positions of training data in their distributions and training dynamics. Experiments across four benchmarks show consistent improvements, achieving average gains of 5.49% over training without selection and 2.89% over the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。