让智能体更懂模糊指令,自动挑出好辨别的对比项
CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
- 用对比学习构建轨迹嵌入空间,拉大相似片段距离
- 在真实人类反馈下性能超越基线方法,提升标注效率
- 适合需要高效对齐人类意图的交互式强化学习场景
基于偏好强化学习(PbRL)通过人类偏好比较推断奖励函数,避免了显式奖励设计,更好地对齐人类意图。然而,面对相似片段时人类常难以给出明确偏好,导致标注效率低,限制了PbRL的实际应用。为此,我们提出一种离线PbRL方法:对比学习以解决模糊反馈(CLARIFY),其学习包含偏好信息的轨迹嵌入空间,使区分度高的片段在空间中相距更远,从而选出更清晰的对比查询。大量实验表明,CLARIFY在非理想教师与真实人类反馈设置下均优于基线。该方法不仅能选择更具区分性的查询,还能学习有意义的轨迹嵌入。
原文摘要 · Abstract (English)
Preference-based reinforcement learning (PbRL) bypasses explicit reward engineering by inferring reward functions from human preference comparisons, enabling better alignment with human intentions. However, humans often struggle to label a clear preference between similar segments, reducing label efficiency and limiting PbRL's real-world applicability. To address this, we propose an offline PbRL method: Contrastive LeArning for ResolvIng Ambiguous Feedback (CLARIFY), which learns a trajectory embedding space that incorporates preference information, ensuring clearly distinguished segments are spaced apart, thus facilitating the selection of more unambiguous queries. Extensive experiments demonstrate that CLARIFY outperforms baselines in both non-ideal teachers and real human feedback settings. Our approach not only selects more distinguished queries but also learns meaningful trajectory embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。