让视觉运动策略在偏离专家轨迹时仍能稳定运行
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- 用隐空间构建安全边界,区分正常与异常状态
- 仅需少量专家数据即可实现鲁棒操作,无需额外人工标注
- 适合数据稀缺但需高可靠性的机器人操控场景
通过行为克隆训练的视觉运动策略容易受协变量偏移影响,微小轨迹偏差会累积导致失败。现有缓解方法如人机协作修正或合成数据增强,往往成本高、依赖强假设或降低模仿质量。我们提出潜空间策略屏障(Latent Policy Barrier, LPB),受控制屏障函数启发,将专家示范的隐空间嵌入视为隐式屏障,分隔安全的分布内状态与不安全的分布外状态。该方法将精确模仿与分布外恢复解耦:基线扩散策略仅基于专家数据,动力学模型则联合训练专家数据与次优策略回放数据。推理时,动力学模型预测未来隐状态并优化其保持在专家分布内。仿真与真实世界实验均表明,LPB提升了策略鲁棒性与数据效率,在有限专家数据下实现可靠操作,且无需额外人工干预或标注。
原文摘要 · Abstract (English)
Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training distribution through human-in-the-loop corrections or synthetic data augmentation. However, these approaches are often labor-intensive, rely on strong task assumptions, or compromise the quality of imitation. We introduce Latent Policy Barrier, a framework for robust visuomotor policy learning. Inspired by Control Barrier Functions, LPB treats the latent embeddings of expert demonstrations as an implicit barrier separating safe, in-distribution states from unsafe, out-of-distribution (OOD) ones. Our approach decouples the role of precise expert imitation and OOD recovery into two separate modules: a base diffusion policy solely on expert data, and a dynamics model trained on both expert and suboptimal policy rollout data. At inference time, the dynamics model predicts future latent states and optimizes them to stay within the expert distribution. Both simulated and real-world experiments show that LPB improves both policy robustness and data efficiency, enabling reliable manipulation from limited expert data and without additional human correction or annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。