arXiv:2605.28179cs.CL2026-05

提出新方法评估大模型能力,训练中就能判断性能好坏。

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

论文配图:SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
图 1 · 摘自论文原文
  • 用跨任务共性概念生成多样验证数据,避免评测偏差
  • 在17个基准上验证,损失与实际表现高度相关
  • 无需额外评测,训练时即可用于选模型和调参

扩展的缩放定律可预测下游性能,但现有方法存在两大局限:仅关注基准级表现会引入场景特异性偏差,依赖同分布验证损失无法追踪分布变化下的能力提升。本文主张在能力层面研究下游缩放,以捕捉跨任务的共性技能并剔除基准特异性噪声。提出SuperValid框架,通过提炼能力域内基准的核心概念,并扩展为多样化、知识丰富的文本,生成跨分布、能力对齐的验证数据。在17个基准(分属6个能力域)上的实验表明,SuperValid损失与下游性能具有强且稳定的关联性,适用于不同架构、规模及训练数据分布的模型。该方法无需额外评测,可在训练过程中实时计算,支持模型选择、早停和缩放决策。

原文摘要 · Abstract (English)

Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream benchmark performance. However, prior approaches face generalization limitations from two aspects: focusing on benchmark-level performance introduces scenario-specific artifacts, while relying on IID validation loss fails to track capability improvements when training distributions vary. In this work, we argue that downstream scaling should be studied at the capability level, which captures shared skill factors across related tasks while abstracting away benchmark-specific noise. We propose SuperValid, a framework that synthesizes OOD (out-of-distribution), capability-aligned validation data by distilling core concepts from benchmarks within a capability domain and expanding them into diverse, knowledge-rich texts. Extensive experiments spanning 17 benchmarks grouped into 6 capability domains show that SuperValid loss exhibits strong and stable correlation with downstream performance across models of different architectures, scales, and training data distributions. As a training-free metric computable during training without benchmark evaluation, SuperValid enables effective model selection, early stopping, and scaling decisions.

大模型评估能力对齐训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。