无需额外交互,通过反馈动态融合策略实现用户偏好对齐。
Dynamic Policy Fusion for User Alignment Without Re-Interaction
- 基于轨迹级反馈,动态融合用户偏好与预训练策略。
- 在多个环境中同时完成任务目标与用户偏好,效果稳定。
- 零样本适配,避免重训,适合个性化智能体部署场景。
深度强化学习策略虽能最大化任务奖励,却未必符合人类用户的个人偏好。传统方法需重新训练以融入用户特定偏好,但偏好奖励函数通常不可得,重训成本高昂。本文提出一种更实用的方法:利用人类反馈,在不增加环境交互的前提下,将已训练策略与用户意图动态融合。通过分析任务训练所用轨迹的反馈,构建理论支持的动态策略融合机制。实验表明,该方法在多个环境中均能同时达成任务目标并满足用户特定需求,且无需额外交互,属于零样本适配。该方法有效解决了个性化对齐中的效率与可行性问题。
原文摘要 · Abstract (English)
Deep reinforcement learning (RL) policies, although optimal in terms of task rewards, may not align with the personal preferences of human users. To ensure this alignment, a naive solution would be to retrain the agent using a reward function that encodes the user's specific preferences. However, such a reward function is typically not readily available, and as such, retraining the agent from scratch can be prohibitively expensive. We propose a more practical approach - to adapt the already trained policy to user-specific needs with the help of human feedback. To this end, we infer the user's intent through trajectory-level feedback and combine it with the trained task policy via a theoretically grounded dynamic policy fusion approach. As our approach collects human feedback on the very same trajectories used to learn the task policy, it does not require any additional interactions with the environment, making it a zero-shot approach. We empirically demonstrate in a number of environments that our proposed dynamic policy fusion approach consistently achieves the intended task while simultaneously adhering to user-specific needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。