arXiv:2503.00812cs.LG2025-03ACL被引 5

提出BOSE方法,让大模型预训练阶段的评估更稳定可靠。

BOSE: A Systematic Evaluation Method Optimized for Base Models

  • 用填空式提示和新指标替代传统困惑度,减少早期训练评估波动
  • 实验显示评估稳定性提升,基座与指令模型表现一致性显著增强
  • 适合关注模型训练引导、评估可信度的研究者

本文指出基座模型(无后训练)评估存在两大问题:一是预训练初期模型能力不足导致评估结果不稳定,难以指导关键实验如数据消融和缩放定律研究;二是基座模型与对应指令模型之间存在性能差距,难以判断基座模型优劣是否真正带来指令模型提升。为此,我们提出基座模型导向的系统评估方法BOSE,包含两项创新:针对开放任务设计的上下文轻指令提示(ICLiP),针对多选任务的空白困惑度(Blank-ppl),将标准困惑度转为填空形式以缓解早期评估波动;首次引入肯德尔等级相关系数量化评估稳定性和一致性。实验表明,BOSE显著提升了预训练过程中的评估稳定性,并增强了基座与指令模型间的评估一致性,为大模型训练提供了更可靠的指导。

原文摘要 · Abstract (English)

This paper poses two critical issues in evaluating base models (without post-training): (1) Unstable evaluation during training: in the early stages of pre-training, the models lack the capability to answer questions as required, leading to unstable evaluation results. This instability makes it difficult to provide solid conclusions to guide the training, especially for key experiments such as data ablation and scaling law. (2) Inconsistency between base and instruct models: base models generally exhibit poorer evaluation performance compared to corresponding instruct models. This gap poses a challenge for assessing whether a base model with better evaluation can truly lead to a better instruct model. To address these issues, we propose Base model Oriented Systematic Evaluation (BOSE), a method specifically designed to optimize the evaluation of base models. Specifically, BOSE introduces two key innovations: In-Context Light-instruction Prompt (ICLiP) for open-ended tasks and Blank-ppl for multi-choice tasks with candidate options, which transforms the standard perplexity (ppl) metric into a fill-in-the-blank format to mitigate early-stage evaluation fluctuations. Furthermore, we are the first to propose Kendall's rank correlation to quantitatively measure the evaluation stability and consistency. Experimental results demonstrate that BOSE significantly enhances both the stability of evaluations during pre-training and the consistency between base and instruct models, thereby providing more reliable guidance for the LLMs' training.

模型评估大模型训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。