arXiv:2605.16043cs.ROcs.AI2026-05中稿 · ICRA

用物理状态替代视觉图像,让机器人更高效学会解绳结。

Learning Sim-Grounded Policies for Bimanual Rope Manipulation from Human Teleoperation Data

论文配图:Learning Sim-Grounded Policies for Bimanual Rope Manipulation from Human Teleoperation Data
图 1 · 摘自论文原文
  • 用3D粒子状态代替摄像头画面作为输入,提升动作预测精度。
  • 在未见过的绳结配置上,状态模型比视觉模型误差降低30.8%。
  • 适合做少样本下柔性物体操控的机器人学习研究者参考。

绳索等可变形线性物体在家庭和工业场景中广泛存在,但因其无限维配置空间和频繁自遮挡而难以操控。通过人类远程操作数据进行模仿学习是实现双臂操控的有效路径,但其可扩展性受限于人力投入,因此观察空间的选择对小数据集下的泛化至关重要。本研究探讨了视角视觉策略在解结任务中泛化能力不足是否源于观察空间本身,而非策略结构或数据规模。我们在同一套双臂远程操作数据上对比两种基于Transformer的动作分块策略:一种依赖手腕相机的双视角RGB流的视觉策略,另一种则基于通过多视角融合获取的3D粒子状态,并在基于粒子的扩展位置动力学模拟中演化该状态。在开放环路评估中,针对未见过的绳结配置,状态策略在预测初始抓握与拉拽动作时,相比视觉策略降低了30.8%的L1误差,量化了像素与物理一致状态之间的可观测性差距,表明以物理状态为输入能更高效地实现柔性物体操控的机器人学习。

原文摘要 · Abstract (English)

Deformable Linear Objects (DLOs) such as ropes and cables are widely encountered in both household and industrial applications, yet remain challenging to manipulate due to their infinite-dimensional configuration space and frequent self-occlusion. Imitation learning from teleoperation offers a practical path to bimanual DLO manipulation, but its scalability is limited by human effort, making the choice of observation space critical for generalization from small datasets. In this study, we investigate whether the lack of generalization in egocentric visual policies for the knot-untangling task stems from the observation space itself, rather than from the policy architecture or data scale. We compare two Action Chunking with Transformers policies trained on the same bimanual teleoperation data: a vision-based policy conditioned on two egocentric RGB streams from wrist-mounted cameras, and a state-based policy conditioned on the DLO's 3D particle state, extracted from an initial observation via multi-view fusion and evolved in a particle-based eXtended Position-Based Dynamics simulation. Evaluated open-loop on an unseen rope configuration, the state-based policy outperforms its visual counterpart with a 30.8% reduction in L1 error when predicting the initial grasp-and-pull action, quantifying the observability gap between pixels and physics-consistent state, and pointing toward more data-efficient robot learning for the DLO manipulation task from limited human demonstrations.

柔性物体模仿学习状态估计双臂操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。