arXiv:2506.05967cs.AIcs.LG2025-06ICML被引 7

用因果视角解决大模型对齐中的偏好学习难题

Preference Learning for AI Alignment: a Causal Perspective

  • 从因果推断出发,识别偏好学习的三大挑战
  • 揭示传统方法在新提示下的泛化失败机制
  • 适合关注模型对齐与数据偏差的研究者

从偏好数据中进行奖励建模是将大型语言模型(LLMs)与人类价值观对齐的关键步骤,要求模型能稳健泛化到新的提示-响应对。本文提出以因果范式重构该问题,引入因果推断工具箱,识别出持续存在的挑战,如因果误判、偏好异质性以及用户特异性因素导致的混淆。基于因果推断文献,我们明确了可靠泛化的关键假设,并与常见的数据收集实践进行对比。通过展示朴素奖励模型的失效模式,证明了因果启发方法可提升模型鲁棒性。最后,我们提出未来研究与实践的期望目标,倡导针对观测数据固有局限性的针对性干预。

原文摘要 · Abstract (English)

Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs. In this work, we propose to frame this problem in a causal paradigm, providing the rich toolbox of causality to identify the persistent challenges, such as causal misidentification, preference heterogeneity, and confounding due to user-specific factors. Inheriting from the literature of causal inference, we identify key assumptions necessary for reliable generalisation and contrast them with common data collection practices. We illustrate failure modes of naive reward models and demonstrate how causally-inspired approaches can improve model robustness. Finally, we outline desiderata for future research and practices, advocating targeted interventions to address inherent limitations of observational data.

模型对齐因果推断偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。