arXiv:2605.09404cs.LGcs.CL2026-05

用目标验证数据生成参考路径,自动选最优训练样本。

Let the Target Select for Itself: Data Selection via Target-Aligned Paths

  • 基于目标验证集的短时预热流生成参考路径,避免传统方法偏差。
  • 无需梯度或海森矩阵近似,仅通过终点损失下降率评分,效率更高。
  • 路径可复用,适合多数据池筛选,节省计算与存储资源。

目标数据选择旨在从大规模候选样本中挑选能提升特定下游任务性能的样本。现有方法通常通过候选集诱导的轨迹上局部贡献分数的聚合来估计样本价值,但当候选集异质时,参考轨迹可能与目标选子集的动力学不一致,导致参考路径偏差。本文提出一种替代参考路径:在可用的目标验证代理上进行短时、容量受限的预热,得到验证诱导的流动。沿此路径,候选样本通过归一化的终点损失下降率评分,形成无需候选梯度或海森矩阵近似的零阶选择规则。在受控的逻辑回归、视觉和指令微调实验中,该评分与强动态归因基线相当,同时显著降低预热时间和存储开销。此外,由于参考轨迹与特定候选集解耦,同一紧凑预热可跨多个候选池复用,无需重新计算轨迹。

原文摘要 · Abstract (English)

Targeted data selection aims to identify training samples from a large candidate pool that improve performance on a specific downstream task. Many recent methods estimate candidate utility by aggregating local attribution scores along a trajectory induced by the candidate pool. When the pool is heterogeneous, however, this reference trajectory may be misaligned with the dynamics of a target-aligned selected subset, creating what we call reference path bias. We propose an alternative reference path: a validation-induced flow obtained from a short, capacity-limited warmup on the available target validation proxy. Along this path, candidates are scored by a normalized endpoint loss drop, yielding a simple zero-order selection rule that requires no candidate gradients or Hessian approximations. Across controlled logistic, vision, and instruction-tuning experiments, this score is competitive with strong dynamic attribution baselines while substantially reducing warmup and storage cost. Moreover, since the reference trajectory is decoupled from any specific candidate pool, the same compact warmup can be reused across additional pools without recomputing the trajectory.

数据选择目标对齐高效训练零阶方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。