动态调整样本权重,让奖励模型少依赖表面线索,多关注真实质量。
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity
- 通过语义保持的反事实扰动,实时检测样本对表面线索的敏感度。
- 在训练中动态降低高敏感样本权重,提升模型对真实偏好信号的依赖。
- 适用于需要可靠偏好建模的场景,如大模型对齐与评估。
从成对偏好训练的奖励模型常利用表面线索而非真正响应质量。我们提出 DynaCF,一种动态重加权框架,用于缓解奖励模型训练中的捷径学习问题。不同于静态捷径启发式方法,DynaCF 在优化过程中通过施加语义保持的反事实扰动,实时测量捷径敏感性,并追踪当前模型下的边际变化与偏好反转。对捷径敏感度更高的样本在 Bradley-Terry 目标中被动态降权,促使模型减少对表面模式的依赖,转而更多关注任务相关的偏好信号。大量实验表明,DynaCF 能持续提升偏好建模的鲁棒性。
原文摘要 · Abstract (English)
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training. Unlike static shortcut heuristics, DynaCF measures shortcut sensitivity online during optimization by applying semantics-preserving counterfactual perturbations and tracking the resulting margin shifts and preference flips under the current model. Samples with higher shortcut sensitivity are dynamically downweighted in the Bradley-Terry objective, encouraging the model to rely less on superficial patterns and more on task-relevant preference signals. Extensive experiments show that DynaCF consistently improves robustness in preference modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。