arXiv:2410.03717cs.CLcs.AI2024-10被引 13

后训练能显著提升模型能力,而非仅调整风格。

Revisiting the Superficial Alignment Hypothesis

  • 后训练性能随微调样本数呈幂律增长,类似预训练规律。
  • 少量样本仅改风格,大量样本才能提升数学与多跳推理能力。
  • 适合评估模型真实能力的研究者,尤其关注基准测试者。

表面对齐假说认为语言模型的全部能力均在预训练中获得,后训练仅用于调整风格和格式。本文通过实证研究后训练的缩放行为,使用多种尺寸的Llama-3、Mistral和Llama-2模型,在数学推理、编程、指令遵循和多跳推理等任务上,考察了微调样本数量对性能的影响。实验发现,后训练性能随微调样本数呈幂律增长,且该规律覆盖广泛能力。对于数学和多跳推理任务,少量样本仅实现风格对齐,性能未饱和;只有增加样本量,模型推理能力才显著提升。这表明,模型可通过后训练整合新知识,尤其在多跳问答等任务上表现增强。因此,表面对齐假说过于简化,需依赖客观基准而非仅人类偏好对齐来评估模型真实能力。

原文摘要 · Abstract (English)

The Superficial Alignment Hypothesis posits that almost all of a language model's abilities and knowledge are learned during pre-training, while post-training is about giving a model the right style and format. We re-examine these claims by empirically studying the scaling behavior of post-training with increasing finetuning examples and evaluating them using objective task-specific standardized benchmarks. Through experiments with the Llama-3, Mistral, and Llama-2 model families of multiple sizes, we observe that, similar to the pre-training scaling laws, post-training task performance scales as a power law against the number of finetuning examples. This power law relationship holds across a broad array of capabilities, including mathematical reasoning, coding, instruction following, and multihop-reasoning. In addition, for tasks like math and multihop reasoning, we observe that a handful of examples merely align the model stylistically but do not saturate performance on the benchmarks. Model performance is instead correlated with its reasoning ability and it improves significantly with more examples, illustrating the need for holistic evaluation programs leveraging objective benchmarks in addition to measurement of alignment to human preferences. We also observe that language models are not necessarily limited to using knowledge learned during pre-training. With appropriate post-training, a model's ability to integrate new knowledge greatly improves on downstream tasks like multihop question-answering. Taken together, these results shed new light on the Superficial Alignment Hypothesis, suggesting that it is, at best, an over-simplification.

语言模型后训练能力评估幂律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。