arXiv:2606.06892cs.LG2026-06

提出GRASP方法,让数据筛选更懂数据间关系,提升预训练效率。

GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

  • 基于几何惩罚建模数据子集的交互关系,突破传统孤立评分局限。
  • 在大规模数据集上实现更高任务相关性,相关性提升超一倍,成本降近九成。
  • 适用于语言和视觉模型的数据筛选,适合大规模预训练优化场景。

现有可扩展数据归因方法通常为单个训练样本分配独立效用分值,这种加性假设无法捕捉数据冗余与互补覆盖等关键子集动态。本文将归因重新定义为子集级反事实效用预测,提出交互感知的代理模型GRASP。基于理论平滑性下界,GRASP通过二次几何惩罚显式建模子集间交互。为实现预训练规模的高效性且无需隐藏的调优机制,采用低维特征草图结合严格有限置信下界选择协议。大量子集重训练评估表明,GRASP显著优于现有可扩展基线:反事实子集保真度的任务级相关性提升超过一倍,前期伪影构建成本降低近一个数量级。下游诊断进一步显示,该评分机制可迁移至语言模型筛选与跨域视觉选择,为优化大规模预训练语料库提供了稳健基础。

原文摘要 · Abstract (English)

Scalable data attribution methods typically assign isolated utility scores to individual training examples. This prevalent additive assumption fundamentally fails to capture critical subset dynamics, including data redundancy and complementary coverage. In this work, we reframe attribution as subset-level counterfactual utility prediction and introduce GRASP, an interaction-aware surrogate. Grounded in a theoretical smoothness lower bound, GRASP explicitly models subset interactions through a quadratic geometric penalty. To achieve pretraining-scale efficiency without relying on hidden oracle tuning, we couple low-dimensional feature sketches with a strictly finite lower-confidence bound selection protocol. Extensive subset-retraining evaluations demonstrate that GRASP decisively outperforms existing scalable baselines. It more than doubles the task-level rank correlation for counterfactual subset fidelity while reducing upfront artifact construction costs by nearly an order of magnitude. Downstream diagnostics further show that this scoring mechanism transfers to language model curation and cross-domain vision selection, establishing a robust foundation for optimizing massive pretraining corpora.

数据归因预训练几何建模高效筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。