arXiv:2410.14001cs.LGcs.CL2024-10被引 11

用上下文学习实现大模型个性化适配,高效且省算力。

Personalized Adaptation via In-Context Preference Learning

  • 通过上下文学习动态适应用户偏好
  • 在上下文老虎机中表现优于现有方法
  • 离线训练+在线自适应,计算成本低

强化学习从人类反馈(RLHF)广泛用于对齐语言模型与人类偏好。然而,现有方法常忽视个体用户偏好,导致个性化效果不佳。本文提出偏好预训练变压器(PPT),一种利用在线用户反馈实现自适应个性化的新型方法。PPT利用Transformer的上下文学习能力,动态适应个体偏好。该方法分为两个阶段:(1) 离线阶段,使用依赖历史的损失函数训练单一策略模型;(2) 在线阶段,通过上下文学习实现用户偏好适配。我们在上下文老虎机设置中验证了PPT的有效性,结果表明其个性化适配性能优于现有方法,同时显著降低计算开销。研究提示,上下文学习在大语言模型的可扩展、高效个性化中具有潜力。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is widely used to align Language Models (LMs) with human preferences. However, existing approaches often neglect individual user preferences, leading to suboptimal personalization. We present the Preference Pretrained Transformer (PPT), a novel approach for adaptive personalization using online user feedback. PPT leverages the in-context learning capabilities of transformers to dynamically adapt to individual preferences. Our approach consists of two phases: (1) an offline phase where we train a single policy model using a history-dependent loss function, and (2) an online phase where the model adapts to user preferences through in-context learning. We demonstrate PPT's effectiveness in a contextual bandit setting, showing that it achieves personalized adaptation superior to existing methods while significantly reducing the computational costs. Our results suggest the potential of in-context learning for scalable and efficient personalization in large language models.

个性化上下文学习大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。