arXiv:2409.18121cs.ROcs.CV2024-09CoRL被引 71

机器人通过看单视角视频学习操作可动物体,无需训练即可模仿动作。

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

论文配图:Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
图 1 · 摘自论文原文
  • 用可微分渲染重建单视频中物体部件的4D运动轨迹
  • 在9个物体上实现60%端到端成功率,每阶段达87%成功率
  • 适合无标注数据、无任务训练的机器人动作模仿场景

人类可通过观察他人学习操作新物体;若让机器人具备此类学习能力,将实现自然的行为指令接口。本文提出「机器人看机器人做」(RSRD)方法,仅需一个单目RGB人类示范视频和一个静态多视角物体扫描,即可实现对可动物体的操作模仿。首先提出4D可微分部件模型(4D-DPM),通过可微分渲染从单视频恢复3D部件运动,采用部件中心特征场进行迭代优化,并引入几何正则化,仅凭单一视频即可恢复3D运动。基于此4D重建,机器人规划双臂动作以复现物体部件运动轨迹。通过部件中心轨迹表示,RSRD聚焦于复制演示的意图行为,同时考虑机器人自身形态限制,而非模仿手部运动。在9个物体上各进行10次试验,共90次测试中,4D-DPM在真实标注的3D部件轨迹上实现高精度跟踪,整个系统端到端成功率达60%,各阶段平均成功率87%。值得注意的是,该方法仅使用大型预训练视觉模型提取的特征场,无需任何任务特定训练、微调、数据收集或标注。

原文摘要 · Abstract (English)

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation from a single monocular RGB human demonstration given a single static multi-view object scan. We first propose 4D Differentiable Part Models (4D-DPM), a method for recovering 3D part motion from a monocular video with differentiable rendering. This analysis-by-synthesis approach uses part-centric feature fields in an iterative optimization which enables the use of geometric regularizers to recover 3D motions from only a single video. Given this 4D reconstruction, the robot replicates object trajectories by planning bimanual arm motions that induce the demonstrated object part motion. By representing demonstrations as part-centric trajectories, RSRD focuses on replicating the demonstration's intended behavior while considering the robot's own morphological limits, rather than attempting to reproduce the hand's motion. We evaluate 4D-DPM's 3D tracking accuracy on ground truth annotated 3D part trajectories and RSRD's physical execution performance on 9 objects across 10 trials each on a bimanual YuMi robot. Each phase of RSRD achieves an average of 87% success rate, for a total end-to-end success rate of 60% across 90 trials. Notably, this is accomplished using only feature fields distilled from large pretrained vision models -- without any task-specific training, fine-tuning, dataset collection, or annotation. Project page: https://robot-see-robot-do.github.io

动作模仿4D重建单目视觉机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。