弱监督训练的强模型在分布偏移下表现崩溃,新方法提升跨数据集泛化能力。
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

- 用零样本分布偏移测试弱到强奖励模型迁移能力
- 发现强模型在原分布下表现好但跨数据集性能骤降
- 提出表示锚定机制,兼顾迁移性与原任务表现
弱到强(W2S)泛化是可扩展监督的有前景框架,但现有评估多在匹配训练-测试分布下进行。本文研究零样本分布偏移下的W2S偏好学习,发现基于弱偏好标签训练的强模型虽在原分布内表现良好,却无法跨偏好数据集迁移。我们揭示一种表征失败模式:弱监督微调会将强模型拉向源域特征,而非保持通用偏好表征。为此,提出表示锚定(Anchor),一种简单有效的正则化方法,在微调中约束对预训练强模型表征空间的过度偏离,同时允许任务相关适应。在多个偏好领域、数据集和模型家族上,Anchor 均显著提升分布外迁移性能,且维持竞争力的分布内表现。本工作提供评估协议、迁移感知指标与方法,揭示当前W2S奖励建模的隐性脆弱性,并为更鲁棒的偏好迁移提供可行路径。
原文摘要 · Abstract (English)
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore, we study W2S preference learning under zero-shot distribution shift and find that strong students trained on weak preference labels can appear successful in-distribution while failing to transfer across preference datasets. We provide evidence for a representational failure mode in which weak-supervised fine-tuning can pull the strong model toward source-domain features instead of maintaining broadly transferable preference representations. To mitigate this, we propose Representation Anchoring (Anchor), a simple yet effective regularizer that constrains excessive drift from the pretrained strong model's representation space during fine-tuning, while still allowing task-relevant adaptation. Across preference domains, datasets, and model families, Anchor consistently improves out-of-distribution transfer while maintaining competitive in-distribution performance. Together, our evaluation protocol, transfer-aware metrics, and method expose hidden brittleness in current W2S reward modeling and provide a practical path toward more robust preference transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。