arXiv:2605.12908stat.MLcs.LG2026-05

弱模型输出指导强模型学习新任务,同时保留原有能力。

The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge

论文配图:The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
图 1 · 摘自论文原文
  • 用弱模型的输出监督强模型微调,实现特征提取。
  • 微调后强模型学会目标任务,且未遗忘其他预训练能力。
  • 适合研究大模型对齐与持续学习的科研人员。

弱到强(W2S)泛化指用一个较弱、任务专精的模型输出来微调一个更强的模型,被提出作为对齐超人级人工智能系统的一种方法。现有理论分析或固定学生模型表示,或局限于特定场景。多步随机梯度下降能否在保留多样预训练能力的同时实现特征学习尚不明确。本文在两层神经网络的奖励模型学习设定下研究W2S。强模型具有组织在低维子空间 $V_k$ 中的预训练表示,并在任务 $κ$ 上由弱模型监督微调。我们证明强模型能高效学习任务 $κ$,在激发预训练知识的同时保持通用能力。这确立了特征学习范式下的W2S泛化:强模型通过W2S训练获得目标特征方向,而非预先给定。此外,W2S保留了预训练的非目标特征,而标准监督微调在非目标特征方向与目标相关时会导致灾难性遗忘。合成数据上的数值实验验证了理论结果。

原文摘要 · Abstract (English)

Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $κ$. We prove that the strong model efficiently learns task $κ$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.

模型对齐特征学习持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。