用人类语义标签引导机器人发现更安全、多样且有用的技能。
Leveraging Human Feedback for Semantically-Relevant Skill Discovery

- 通过人类标注语义标签,让机器学习更符合人类认知的行为模式。
- 在2D导航和4个运动环境中,显著提升技能的语义多样性和相关性。
- 适合需要安全、可控、贴近人类意图的强化学习应用。
无监督强化学习中的技能发现旨在让智能体自主探索多样且有用的行為。然而,不受约束的方法可能产生不安全、不道德或与人类目标不符的行为。为降低风险并提升实用性,近期研究利用人类偏好反馈来引导发现过程。但这类方法反馈效率低,且难以应对包含跑步、跳跃、行走等多种不同技能的复杂技能空间。为此,我们提出语义标注(semantic labelling)——一种新型且高效利用人类反馈的方法,借助人类认知优势识别并标注语义上有意义的行为。基于此,我们提出语义相关技能发现(SRSD),一种人机协作方法:收集人类对行为的语义标签,并据此学习奖励函数,以促进技能在语义上的多样性与相关性。在2D导航环境及四个运动环境中进行实验表明,SRSD能有效提升语义多样性,同时可扩展至大量不同类型的行为。
原文摘要 · Abstract (English)
Unsupervised skill discovery in reinforcement learning aims to intrinsically motivate agents to discover diverse and useful behaviours. However, unconstrained approaches can produce unsafe, unethical, or misaligned behaviours. To mitigate these risks and improve the practical desireability of discovered skills, recent work grounds the discovery process by leveraging human preference feedback. However, preference-based approaches are feedback-inefficient and inherently ill-equipped to deal with skill spaces composed of a variety of different skills such as running, jumping, walking, etc. To overcome this limitation, we introduce semantic labelling, a novel and feedback-efficient approach that leverages human cognitive strengths to identify and label semantically meaningful behaviours. Based on semantic labelling, we propose Semantically Relevant Skill Discovery (SRSD), a novel human-in-the-loop approach that collects semantic labels from human feedback and learns a reward function to encourage skills to be more semantically diverse and relevant. Through our experiments in a 2D navigation environment and four locomotion environments, we demonstrate that SRSD can improve semantic diversity and discover relevant behaviours while scaling effectively to a large variety of behaviours.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。