arXiv:2509.22578cs.RO2025-09

让机器人在新视角下也能模仿动作,提升操作泛化能力。

EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation

  • 通过分割动作、几何变换和逆运动学滤波,实现动作轨迹的视角迁移。
  • 在仿真和真实机器人上分别提升24.6%和16.0%的成功率。
  • 无需多视角数据,自监督生成逼真新视角视觉观察,适合实际部署。

基于视觉-动作策略的模仿学习在机器人操作中表现优异,但对第一人称视角变化敏感。与第三人称视角仅改变相机位置不同,第一人称视角同时改变相机姿态和机器人动作坐标系,需联合迁移动作轨迹并合成新视角下的观测。为此,我们提出EgoDemoGen框架,包含两个关键组件:1)EgoTrajTransfer,通过动作技能分割、几何感知变换和逆运动学滤波,将机器人轨迹迁移到新第一人称坐标系;2)EgoViewTransfer,一种条件视频生成模型,融合新视角重投影场景视频与从迁移轨迹渲染的机器人运动视频,合成逼真观测,采用自监督双重投影策略训练,无需多视角数据。仿真与真实场景实验表明,EgoDemoGen在标准与新视角下均显著提升策略成功率,仿真绝对提升24.6%和16.9%,真实机器人提升16.0%和23.0%。此外,EgoViewTransfer在新视角观测生成质量上表现优异。

原文摘要 · Abstract (English)

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-person viewpoint changes that only move the camera, egocentric shifts simultaneously alter both the camera pose and the robot action coordinate frame, making it necessary to jointly transfer action trajectories and synthesize corresponding observations under novel egocentric viewpoints. To address this challenge, we present EgoDemoGen, a framework that generates paired observation--action demonstrations under novel egocentric viewpoints through two key components: 1{)} EgoTrajTransfer, which transfers robot trajectories to the novel egocentric coordinate frame through motion-skill segmentation, geometry-aware transformation, and inverse kinematics filtering; and 2{)} EgoViewTransfer, a conditional video generation model that fuses a novel-viewpoint reprojected scene video and a robot motion video rendered from the transferred trajectory to synthesize photorealistic observations, trained with a self-supervised double reprojection strategy without requiring multi-viewpoint data. Experiments in simulation and real-world settings show that EgoDemoGen consistently improves policy success rates under both standard and novel egocentric viewpoints, with absolute gains of +24.6\% and +16.9\% in simulation and +16.0\% and +23.0\% on the real robot. Moreover, EgoViewTransfer achieves superior video generation quality for novel egocentric observations.

机器人操作视角泛化模仿学习视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。