arXiv:2412.19436stat.MLcs.LG2024-12被引 10

用低秩结构建模用户差异,让大模型更懂不同人的偏好。

Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback

  • 基于用户上下文的低秩偏好模型,高效处理反馈多样性。
  • 在个性化任务中显著优于基线,且对反馈分布变化更鲁棒。
  • 适合需要个性化对齐的对话系统与推荐场景。

从人类反馈中进行强化学习(RLHF)已成为对齐大语言模型与人类偏好的核心方法。然而,由个体差异和偏好多样导致的人类反馈异质性,给奖励学习带来挑战。为此,我们提出一种低秩上下文强化学习框架(LoCo-RLHF),通过融合上下文信息来更好建模异质反馈,同时保持计算效率。该方法基于上下文偏好模型,利用用户上下文与问答对之间交互的内在低秩结构,缓解特征表示的高维问题。此外,我们提出受悲观离线强化学习启发的‘降维空间中的悲观性’(PRS)策略,以应对反馈分布漂移问题。理论证明该策略相较于现有方法具有更紧的次优性差距。大量实验验证了LoCo-RLHF的有效性,在个性化RLHF设置中表现优异,并展现出对分布变化的强鲁棒性。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) has become a cornerstone for aligning large language models with human preferences. However, the heterogeneity of human feedback, driven by diverse individual contexts and preferences, poses significant challenges for reward learning. To address this, we propose a Low-rank Contextual RLHF (LoCo-RLHF) framework that integrates contextual information to better model heterogeneous feedback while maintaining computational efficiency. Our approach builds on a contextual preference model, leveraging the intrinsic low-rank structure of the interaction between user contexts and query-answer pairs to mitigate the high dimensionality of feature representations. Furthermore, we address the challenge of distributional shifts in feedback through our Pessimism in Reduced Subspace (PRS) policy, inspired by pessimistic offline reinforcement learning techniques. We theoretically demonstrate that our policy achieves a tighter sub-optimality gap compared to existing methods. Extensive experiments validate the effectiveness of LoCo-RLHF, showcasing its superior performance in personalized RLHF settings and its robustness to distribution shifts.

强化学习人类反馈个性化低秩建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。