arXiv:2505.02094cs.LGcs.CV2025-05International Conf…被引 32

从稀疏噪声示范中学习鲁棒通用的交互技能

SkillMimic-V2: Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations

  • 构建轨迹图与状态转移场,自动补全示范间的潜在连接
  • 动态课程采样与历史编码,提升训练稳定性和泛化能力
  • 适合需要从不完整示范中学习复杂交互的机器人研究

我们针对强化学习从交互示范(RLID)中的核心挑战——示范噪声与覆盖不足问题提出解决方案。现有数据采集方法虽提供有价值的交互示范,但常产生稀疏、断裂且含噪的轨迹,无法涵盖全部技能变化与状态转移。关键洞察在于:尽管示范稀疏且含噪,仍存在无限多物理可行的轨迹可自然连接不同示范技能或从邻近状态生成,形成连续的技能变化与转移空间。基于此,我们提出两种数据增强技术:拼接轨迹图(STG)用于发现示范技能间的潜在转移,状态转移场(STF)为示范邻域内任意状态建立唯一连接。为有效利用增强数据,我们设计自适应轨迹采样(ATS)策略实现动态课程生成,并引入历史编码机制支持记忆依赖型技能学习。实验在多样交互任务中验证,本方法在收敛稳定性、泛化能力及恢复鲁棒性方面显著优于当前最优方法。

原文摘要 · Abstract (English)

We address a fundamental challenge in Reinforcement Learning from Interaction Demonstration (RLID): demonstration noise and coverage limitations. While existing data collection approaches provide valuable interaction demonstrations, they often yield sparse, disconnected, and noisy trajectories that fail to capture the full spectrum of possible skill variations and transitions. Our key insight is that despite noisy and sparse demonstrations, there exist infinite physically feasible trajectories that naturally bridge between demonstrated skills or emerge from their neighboring states, forming a continuous space of possible skill variations and transitions. Building upon this insight, we present two data augmentation techniques: a Stitched Trajectory Graph (STG) that discovers potential transitions between demonstration skills, and a State Transition Field (STF) that establishes unique connections for arbitrary states within the demonstration neighborhood. To enable effective RLID with augmented data, we develop an Adaptive Trajectory Sampling (ATS) strategy for dynamic curriculum generation and a historical encoding mechanism for memory-dependent skill learning. Our approach enables robust skill acquisition that significantly generalizes beyond the reference demonstrations. Extensive experiments across diverse interaction tasks demonstrate substantial improvements over state-of-the-art methods in terms of convergence stability, generalization capability, and recovery robustness.

强化学习技能学习数据增强机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。