arXiv:2510.07313cs.CVcs.RO2025-10被引 21

用主视角生成机械臂腕部视角视频,提升机器人操作性能

WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation

论文配图:WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
图 1 · 摘自论文原文
  • 基于锚点视角重建几何一致的腕部视角4D点云
  • 生成时间连贯的腕部视角视频,提升空间一致性
  • 适合需要精细手物交互的机器人视觉模型研究

腕部视角观测对视觉-语言-动作(VLA)模型至关重要,因其能捕捉精细的手物交互,直接提升操作性能。然而大规模数据集很少包含此类记录,导致丰富锚点视角与稀缺腕部视角之间存在显著差距。现有世界模型无法弥合该差距,因它们需以腕部视角首帧为起点,无法仅从锚点视角生成。近期视觉几何模型如VGGT具备几何与跨视角先验,使应对极端视角变化成为可能。受此启发,我们提出WristWorld——首个仅凭锚点视角生成腕部视角视频的4D世界模型。WristWorld分为两阶段:(i) 重建阶段,扩展VGGT并引入空间投影一致性(SPC)损失,估计几何一致的腕部视角姿态与4D点云;(ii) 生成阶段,利用视频生成模型从重建视角合成时间连贯的腕部视角视频。在Droid、Calvin和Franka Panda上的实验表明,该方法达到最先进的视频生成效果,具有优异的空间一致性,同时提升VLA性能,在Calvin上平均任务完成长度提升3.81%,关闭了42.4%的锚点-腕部视角差距。

原文摘要 · Abstract (English)

Wrist-view observations are crucial for VLA models as they capture fine-grained hand-object interactions that directly enhance manipulation performance. Yet large-scale datasets rarely include such recordings, resulting in a substantial gap between abundant anchor views and scarce wrist views. Existing world models cannot bridge this gap, as they require a wrist-view first frame and thus fail to generate wrist-view videos from anchor views alone. Amid this gap, recent visual geometry models such as VGGT emerge with geometric and cross-view priors that make it possible to address extreme viewpoint shifts. Inspired by these insights, we propose WristWorld, the first 4D world model that generates wrist-view videos solely from anchor views. WristWorld operates in two stages: (i) Reconstruction, which extends VGGT and incorporates our Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds; (ii) Generation, which employs our video generation model to synthesize temporally coherent wrist-view videos from the reconstructed perspective. Experiments on Droid, Calvin, and Franka Panda demonstrate state-of-the-art video generation with superior spatial consistency, while also improving VLA performance, raising the average task completion length on Calvin by 3.81% and closing 42.4% of the anchor-wrist view gap.

机器人操作4D建模视角生成视觉-语言-动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。