arXiv:2503.19201cs.LGcs.AI2025-03被引 5

用低秩适配提升个性化强化学习对齐效率

A Shared Low-Rank Adaptation Approach to Personalized RLHF

  • 在所有个性化奖励模型的参数空间中应用低秩适配
  • 仅需少量本地数据即可高效学习个性化奖励模型
  • 适合需要个性化对齐且数据有限的AI系统开发

强化学习从人类反馈(RLHF)已成为对齐人工智能系统与人类价值观的关键技术,在微调大语言模型方面取得显著成功。然而,现有框架通常假设人类偏好相对同质,可由单一统一的奖励模型捕捉,忽略了个体间的内在差异性,限制了在个性化场景下的适应能力,并可能引发误对齐,降低用户满意度与信任。本文将低秩适配(LoRA)引入个性化RLHF框架,在所有个性化奖励函数的聚合参数空间中应用LoRA,从而实现从潜在有限本地数据中高效学习个性化奖励模型。该方法利用本地真实奖励模型间的共享结构,同时支持个体化调整,无需依赖先前工作中的强共享表示假设。我们进一步为该方法建立了样本复杂度保证。理论分析表明,所提方法能有效捕捉异质人类偏好中的共性和个体特异性结构,兼顾个性化需求与实际数据约束。真实数据集上的实验结果验证了算法在个性化RLHF设置中的高效性。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) has emerged as a pivotal technique for aligning artificial intelligence systems with human values, achieving remarkable success in fine-tuning large language models. However, existing RLHF frameworks often assume that human preferences are relatively homogeneous and can be captured by a single, unified reward model. This assumption overlooks the inherent diversity and heterogeneity across individuals, limiting the adaptability of RLHF to personalized scenarios and risking misalignments that can diminish user satisfaction and trust in AI systems. In this paper, we address these challenges by introducing Low-Rank Adaptation (LoRA) into the personalized RLHF framework. We apply LoRA in the the aggregated parameter space of all personalized reward functions, thereby enabling efficient learning of personalized reward models from potentially limited local datasets. Our approach exploits potential shared structures among the local ground-truth reward models while allowing for individual adaptation, without relying on restrictive assumptions about shared representations as in prior works. We further establish sample complexity guarantees for our method. Theoretical analysis demonstrates the effectiveness of the proposed approach in capturing both shared and individual-specific structures within heterogeneous human preferences, addressing the dual challenge of personalization requirements and practical data constraints. Experimental results on real-world datasets corroborate the efficiency of our algorithm in the personalized RLHF setting.

个性化对齐低秩适配强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。