用轻量嵌入适配扩散模型,快速匹配用户偏好。
Latent Embedding Adaptation for Human Preference Alignment in Diffusion Planners
- 用可学习的偏好嵌入直接优化生成轨迹
- 在真实用户偏好上表现优于RLHF和LoRA
- 适合需要快速个性化决策的智能系统
本文针对自动化决策系统中轨迹个性化难题,提出一种资源高效的方法,通过预训练条件扩散模型结合偏好潜变量嵌入(PLE),在无奖励的离线数据集上训练。PLE作为紧凑表示捕捉特定用户偏好。利用所提偏好反演方法直接优化可学习的PLE,实现对人类偏好的更优对齐,显著优于强化学习从人类反馈(RLHF)和低秩适应(LoRA)等现有方案。为贴近实际应用,我们基于真实用户偏好构建了涵盖多样化高回报轨迹的基准实验。
原文摘要 · Abstract (English)
This work addresses the challenge of personalizing trajectories generated in automated decision-making systems by introducing a resource-efficient approach that enables rapid adaptation to individual users' preferences. Our method leverages a pretrained conditional diffusion model with Preference Latent Embeddings (PLE), trained on a large, reward-free offline dataset. The PLE serves as a compact representation for capturing specific user preferences. By adapting the pretrained model using our proposed preference inversion method, which directly optimizes the learnable PLE, we achieve superior alignment with human preferences compared to existing solutions like Reinforcement Learning from Human Feedback (RLHF) and Low-Rank Adaptation (LoRA). To better reflect practical applications, we create a benchmark experiment using real human preferences on diverse, high-reward trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。