arXiv:2506.01987cs.LGcs.AI2025-06

同时优化数据样本和目标,能显著提升模型训练效率

Equally Critical: Samples, Targets, and Their Mappings in Datasets

  • 提出样本与目标映射的新分类体系
  • 实验证明目标质量比样本数量影响更大
  • 适合关注数据高效训练的研究者

数据天然具有样本和目标双重属性。针对目标,知识蒸馏通过教师生成的软目标监督已被广泛用于加速模型收敛;而近期数据高效学习则聚焦样本优化(如数据蒸馏),却忽视了目标的重要性。这种割裂促使我们研究样本与目标如何共同影响训练动态。为此,我们基于样本-目标交互构建现有范式的分类体系,归纳出不同映射策略。在此基础上,提出统一损失框架以评估其对训练效率的影响。在多种策略上开展大量实证研究,系统分析目标与样本类型、数量、质量变化对训练的影响,提炼出六条关键洞见,以提升训练有效性。

原文摘要 · Abstract (English)

Data inherently possesses dual attributes: samples and targets. For targets, knowledge distillation has been widely employed to accelerate model convergence, primarily relying on teacher-generated soft target supervision. Conversely, recent advancements in data-efficient learning have emphasized sample optimization techniques, such as dataset distillation, while neglected the critical role of target. This dichotomy motivates our investigation into understanding how both sample and target collectively influence training dynamic. To address this gap, we first establish a taxonomy of existing paradigms through the lens of sample-target interactions, categorizing them into distinct sample-to-target mapping strategies. Building upon this foundation, we then propose a novel unified loss framework to assess their impact on training efficiency. Through extensive empirical studies on our proposed strategies, we comprehensively analyze how variations in target and sample types, quantities, and qualities influence model training, providing six key insights to enhance training efficacy.

数据蒸馏知识蒸馏训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。