用状态门控专家模型融合手持与远程操作数据,提升接触密集操作成功率。
Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts

- 基于机器人状态动态切换专家模型,分阶段使用不同监督信号。
- 在三个接触密集任务中,成功率比仅用手持数据提升最高36.7%。
- 只需少量远程示范,即可显著改善高刚度接触场景下的安全性。
手持式数据采集系统(如UMI)可高效获取多样环境下的数据,但仅记录观测动作而非机器人控制器的实际目标动作。相比之下,远程操作能直接捕获目标动作,但成本过高。我们发现,在无接触的自由空间阶段,手持轨迹具有有效性;但在接触敏感阶段,高刚度跟踪观测轨迹会产生过大且不安全的接触力。为此,我们研究了两类监督信号在接触密集操作中的协同机制,发现仅对基础手持策略失效的任务段补充少量远程示范,即可实现高效混合训练。然而,简单拼接两类数据会劣化性能。为此提出BRIDGE:一种基于状态门控的扩散策略混合专家模型,根据当前机器人状态路由至特定任务阶段的专家头。该方法在接触敏感段有效利用目标动作,使三类接触密集任务的成功率相较仅用手持数据的基线最高提升36.7%。
原文摘要 · Abstract (English)
Handheld data collection systems, such as the Universal Manipulation Interface (UMI), enable scalable data collection across diverse environments but only capture observed actions rather than the desired actions executed by a robot controller. In contrast, teleoperation captures desired actions directly, but is prohibitively time-consuming to collect. We revisit this trade-off through the lens of action validity across task phases. We observe that handheld trajectories provide valid supervision in tolerant, free-space phases, but lack dynamic feasibility in contact-sensitive phases, where tracking observed trajectories at high stiffness produces large, unsafe contact forces. We study the interaction between these two supervision types for contact-rich manipulation and find that training policies that combine handheld data with a small number of targeted teleoperated demonstrations provide an efficient hybrid strategy. Specifically, rather than teleoperating the entire task, we only collect partial teleoperated demonstrations for task segments where base handheld policies fail. However, naively mixing handheld and teleoperated phase-specific data yields worse performance than training on handheld data alone. To address this mismatch between observed and desired supervision, we propose Bi-modal Routing for Imitation Data via Gated Experts (BRIDGE), a mixture of diffusion policy experts that routes between specialist task phase heads conditioned on the current robot state. Notably, our approach enables task-phase specific use of desired actions during contact sensitive segments and improves success rates over handheld-only baselines by up to 36.7% across three contact-rich manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。