arXiv:2501.17431cs.LGcs.RO2025-01中稿 · the 24th Internati…被引 4

用人类反馈引导强化学习,发现安全且有用的多样化技能

Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment

  • 引入人类反馈优化技能多样性与价值对齐
  • 在2D导航和SafetyGymnasium中实现安全有效的技能发现
  • 支持可配置的多样性-对齐权衡,适合实际应用

无监督强化学习中的技能发现旨在模仿人类自主发现多样化行为的能力。然而,现有方法通常缺乏约束,在复杂环境中常发现不安全或不实用的技能。为此,我们提出人类对齐技能发现(HaSD)框架,通过融入人类反馈来发现更安全、更符合人类价值观的技能。HaSD同时优化技能多样性与人类价值观对齐性,确保整个发现过程始终具备对齐性,避免探索无效技能的低效问题。我们在2D导航与SafetyGymnasium环境中验证了其有效性,结果显示HaSD能发现多样、安全且适用于下游任务的技能。最后,我们进一步扩展了HaSD,实现了可配置的多种技能,支持不同层次的多样性与对齐性权衡,适用于实际场景。

原文摘要 · Abstract (English)

Unsupervised skill discovery in Reinforcement Learning aims to mimic humans' ability to autonomously discover diverse behaviors. However, existing methods are often unconstrained, making it difficult to find useful skills, especially in complex environments, where discovered skills are frequently unsafe or impractical. We address this issue by proposing Human-aligned Skill Discovery (HaSD), a framework that incorporates human feedback to discover safer, more aligned skills. HaSD simultaneously optimises skill diversity and alignment with human values. This approach ensures that alignment is maintained throughout the skill discovery process, eliminating the inefficiencies associated with exploring unaligned skills. We demonstrate its effectiveness in both 2D navigation and SafetyGymnasium environments, showing that HaSD discovers diverse, human-aligned skills that are safe and useful for downstream tasks. Finally, we extend HaSD by learning a range of configurable skills with varying degrees of diversity alignment trade-offs that could be useful in practical scenarios.

强化学习技能发现人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。