将生成视频转化为可执行机器人操作轨迹,确保视觉运动与物理现实一致。
GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency

- 通过稀疏SE(3)模型验证生成视频中2D运动是否与真实场景几何一致。
- 仅传递符合物理约束的运动,结合抓取姿态生成可执行轨迹。
- 适合需要视觉引导但要求高可靠性的机器人操作任务。
生成视频为机器人操作提供了有用的视觉运动先验,但其视觉合理性不等于物理可执行性。生成视频通常缺乏度量几何、抓取定位、机器人运动学可行性及执行时反馈,导致直接重放轨迹在真实世界中不可靠。本文提出GenVid2Robot框架,将生成视频运动转换为可执行的真实机器人操作轨迹。给定初始RGB-D观测和任务指令,该方法从首帧提取任务相关的语义锚点,追踪其在生成视频候选中的运动,并在稀疏相对SE(3)模型下验证2D运动是否能由首帧RGB-D锚点解释。因此,生成视频被视为不确定的视觉运动假设,仅几何一致的运动才传给机器人。接受的相对运动应用于由掩码约束抓取选出的真实抓取时刻工具中心点(TCP)位置,生成与视觉运动先验和物理抓取配置一致的抓取条件执行轨迹。为减少因RGB-D噪声、标定残差及微小接触引起的执行偏差,引入有界深度补偿模块,在不假设全程在线重规划的前提下修正局部深度方向误差。真实机器人实验表明,GenVid2Robot通过稀疏度量几何、抓取约束、机器人可行性检查和有界执行反馈,显著提升了生成视频引导操作的可靠性。
原文摘要 · Abstract (English)
Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic feasibility, and execution-time feedback, which makes direct trajectory replay unreliable in real-world manipulation. This paper presents GenVid2Robot, a rigid-geometric consistency framework that converts generated video motion into executable real-robot manipulation trajectories. Given an initial RGB-D observation and a task instruction, GenVid2Robot samples task-relevant semantic anchors from the real first frame, tracks these anchors through generated video candidates, and verifies whether the resulting 2D motion can be explained by first-frame RGB-D anchors under a sparse relative $SE(3)$ model. In this way, generated videos are treated as uncertain visual motion hypotheses rather than direct robot demonstrations. Only geometrically consistent motion is transferred to the robot. The accepted relative motion is then applied to the real grasp-time TCP pose selected by mask-constrained grasping, producing a grasp-conditioned execution trajectory that is consistent with both the visual motion prior and the physical grasp configuration. To reduce execution mismatch caused by RGB-D noise, calibration residuals, and small contact-induced displacement, a bounded depth-compensation module corrects local depth-direction errors without assuming full online replanning. Real-robot experiments demonstrate that GenVid2Robot improves the reliability of generated-video-guided manipulation by grounding visual motion priors with sparse metric geometry, grasp constraints, robot feasibility checking, and bounded execution feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。