用无监督数据构建语义连贯潜空间,实现少反馈下的安全高效技能发现。
COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space

- 基于密集无监督数据构建语义一致的技能潜空间,结构清晰可引导。
- 仅需少量在线反馈即可学习多样且符合人类意图的技能,避免危险行为。
- 无需额外训练模型,直接生成引导信号,适合真实场景快速部署。
无监督技能发现(USD)旨在不依赖奖励函数的情况下学习多样化行为,但常因随机探索产生无关或危险行为。有监督技能发现(GSD)通过引入人类意图聚焦有意义区域,但现有方法通常需训练额外引导模型,依赖预设规则或专家示范,在稀疏、在线收集的人类反馈下效果不佳。为此,我们提出COLLIE,一种利用密集无监督数据构建语义连贯技能潜空间的GSD框架。该潜空间结构良好,支持以稀疏在线反馈实现可靠引导。其语义一致性使引导信号可免训练构建,无需在技能学习外再训练额外模型。理论分析验证了免训练引导信号的有效性。跨多种状态空间与像素空间任务的实验表明,COLLIE能学习多样化、人类对齐的技能,规避危险行为,并在极少人类反馈下实现更优下游性能。
原文摘要 · Abstract (English)
Unsupervised skill discovery (USD) aims to learn diverse behaviors without reward functions, but often results in task-irrelevant or hazardous behaviors due to uniform exploration. Guided skill discovery (GSD) addresses this issue by incorporating human intent to focus exploration on meaningful regions. However, existing GSD methods typically require training additional guidance models, and rely on pre-defined rules or expert demonstration, which can be ineffective under sparse, online-collected human feedback. To overcome this, we propose COLLIE, a GSD framework that leverages dense unsupervised data to construct a semantically coherent skill latent space. This latent space is well-structured, enabling reliable guidance with sparse online feedback. Moreover, its semantic coherence property enables training-free construction of guidance signals, eliminating the need for additional model training beyond skill learning. Theoretical analysis justifies the effectiveness of our training-free guidance signal, while experiments across diverse state-based and pixel-based tasks show that COLLIE learns diverse, human-aligned skills, avoids hazardous behaviors, and achieves superior downstream performance with minimal human feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。