arXiv:2512.02409cs.LGcs.AI2025-12

揭示数据筛选的极限与加速机制,指出静态删减无效,动态调整才可提速。

Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles

  • 将数据筛选建模为重加权采样分布,通过数据算子特征值分析其影响。
  • 静态删减无法改变谱尾指数,仅能带来有限提升,无法改变模型长期性能。
  • 动态优化可显著加速学习,虽现实系统只能近似,但方向明确有效。

大规模神经网络训练日益依赖数据删减、合成数据生成、跨模型蒸馏、基于人类反馈的强化学习(RLHF)及难度采样等策略。尽管部分方法显著提升训练效率与下游性能,另一些如自动生成的合成数据却常仅增加数据量而未能增强模型能力。本文将数据筛选形式化为重加权采样分布,并将其效果映射至数据诱导算子的特征结构。第一个核心结果表明:静态删减导致有界算子,因此无法改变谱尾指数,仅能实现有限区域内的改进,无法影响神经模型的渐近缩放规律。第二个结果分析时间依赖的数据筛选,证明理想预言机若能追踪谱残差并持续重归一化尾部,可实现可证明的学习加速——尽管实际系统只能近似此行为。

原文摘要 · Abstract (English)

Large-scale neural models are increasingly trained with data pruning, synthetic data generation, cross-model distillation, reinforcement learning from human feedback (RLHF), and difficulty-based sampling. While several of these data-centric strategies reliably improve training efficiency and downstream performance, others fail to provide meaningful gains -- most notably self-generated synthetic data, which often increases dataset volume without enhancing model capability. We formalize data curation as reweighting the sampling distribution and map its effect onto the eigenstructure of the data-induced operator. Our first main result shows that \textbf{static pruning induces a bounded operator and therefore cannot change the spectral tail exponent}; it provides at most finite-region improvements and cannot alter asymptotic neural scaling. Our second result analyzes \textbf{time-dependent data curation}, showing that an ideal oracle capable of tracking spectral residuals and continuously re-normalizing the tail can provably accelerate learning -- although practical systems can only approximate this behavior.

数据筛选谱分析模型训练加速机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。