arXiv:2506.13216cs.CL2025-06ACL被引 2

提出能力显著向量,让损失预测更贴合下游任务表现

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law

  • 将总损失分解并为不同词元分配权重,匹配特定元能力
  • 在多个基准上显著提升下游任务性能预测准确率
  • 适合关注模型性能可预测性与计算资源优化的研究者

扩展定律建立了训练计算量与验证损失之间的关系,使研究者能够有效预测不同计算水平下模型的损失趋势。然而,验证损失与模型下游能力之间仍存在差距,使得扩展定律难以直接用于下游任务性能预测。通常情况下,损失代表预测词元的累积惩罚,隐含假设所有词元重要性相同。但我们的研究发现,在不同训练数据分布下,无法直接建立下游能力与计算量或词元损失之间的关系。为弥合验证损失与下游任务能力之间的差距,本文提出能力显著向量(Capability Salience Vector),将整体损失分解并为词元分配不同重要性权重,以评估特定元能力,从而在模型能力层面实现验证损失与下游任务表现的对齐。在多个主流基准上的实验表明,所提出的向量能显著提升语言模型在下游任务上的性能预测能力。

原文摘要 · Abstract (English)

Scaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different levels of computation. However, a gap still remains between validation loss and the model's downstream capabilities, making it untrivial to apply scaling law to direct performance prediction for downstream tasks. The loss typically represents a cumulative penalty for predicted tokens, which are implicitly considered to have equal importance. Nevertheless, our studies have shown evidence that when considering different training data distributions, we cannot directly model the relationship between downstream capability and computation or token loss. To bridge the gap between validation loss and downstream task capabilities, in this work, we introduce Capability Salience Vector, which decomposes the overall loss and assigns different importance weights to tokens to assess a specific meta-capability, aligning the validation loss with downstream task performance in terms of the model's capabilities. Experiments on various popular benchmarks demonstrate that our proposed Capability Salience Vector could significantly improve the predictability of language model performance on downstream tasks.

扩展定律语言模型性能预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。