arXiv:2605.11134cs.LGcs.AI2026-05被引 2

发现并解决语言模型偏好学习中的虚假相关问题

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training

论文配图:Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
图 1 · 摘自论文原文
  • 通过理论分析揭示虚假相关形成的两种机制
  • 证明数据越多反而更依赖虚假特征,存在不可消除的脆弱性
  • 提出用等效偏好对增强训练,可针对性抑制虚假学习

当前基于偏好的学习方法(如 DPO)会诱导语言模型依赖虚假相关,导致逢迎行为和长度偏好,未来可能引发目标误泛化。本文针对对数线性策略提供统一理论分析,揭示虚假学习在总体层面通过均值虚假偏差与因果-虚假相关泄漏两条路径产生。进一步表明,这种依赖导致无法通过增加同分布数据来缓解的分布偏移脆弱性。为此,提出“等效训练”策略,利用等效效用偏好对进行数据增强,引入数据驱动正则化。实证显示该方法可选择性降低虚假学习,不损害因果学习。理论与实验验证了其在对数线性模型、神经网络及大语言模型上的有效性。

原文摘要 · Abstract (English)

Preference learning methods like Direct Preference Optimization (DPO) are known to induce reliance on spurious correlations, leading to sycophancy and length bias in today's language models and potentially severe goal misgeneralization in future systems. In this work, we provide a unified theoretical analysis of this phenomenon, characterizing the mechanisms of spurious learning, its consequences on deployment, and a provable mitigation strategy. Focusing on log-linear policies, we show that standard preference-learning objectives induce reliance on spurious features at the population level through two channels: mean spurious bias and causal-spurious correlation leakage. We then show that this reliance creates an irreducible vulnerability to distribution shift: more data from the same training distribution fails to reduce the model's dependence on spurious features. To address this, we propose tie training, a data augmentation strategy using ties (equal-utility preference pairs) to introduce data-driven regularization. We demonstrate that this approach selectively reduces spurious learning without degrading causal learning. Finally, we validate our theory on log-linear models and provide empirical evidence that both the spurious learning mechanisms and the benefits of tie training persist for neural networks and large language models.

偏好学习虚假相关模型鲁棒性正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。