arXiv:2602.03793cs.ROcs.CV2026-02被引 16

用动作掩码让视频生成模型理解机器人世界,统一多形态控制

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

  • 将动作指令转为像素对齐的实体掩码,通过控制网络注入生成模型
  • 在双臂和单臂数据集上提升视频生成质量,支持新视角和复杂场景
  • 适合研究具身智能、机器人视觉生成及跨模态控制的开发者

具身世界模型在机器人领域展现出巨大潜力,通常依赖大规模互联网视频或预训练视频生成模型来丰富视觉与运动先验。然而,现有方法仍存在动作坐标空间与视频像素空间不匹配、对相机视角敏感、不同实体架构不统一等关键挑战。为此,我们提出BridgeV2W,将坐标空间动作转换为从URDF和相机参数渲染出的像素对齐实体掩码,并通过类似ControlNet的路径注入预训练视频生成模型。该方法使动作控制信号与预测视频对齐,加入视点相关条件以适应不同相机视角,并实现跨实体的统一架构。为缓解对静态背景的过拟合,进一步引入基于流的运动损失,聚焦学习动态且任务相关的区域。在单臂(DROID)和双臂(AgiBot-G1)数据集上的实验表明,相比现有最先进方法,BridgeV2W显著提升了视频生成质量,涵盖多种复杂场景与未见视角。我们还验证了其在下游真实任务中的潜力,包括策略评估与目标条件规划。更多结果见项目网站:https://BridgeV2W.github.io。

原文摘要 · Abstract (English)

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still face key challenges: a misalignment between coordinate-space actions and pixel-space videos, sensitivity to camera viewpoint, and non-unified architectures across embodiments. To this end, we present BridgeV2W, which converts coordinate-space actions into pixel-aligned embodiment masks rendered from the URDF and camera parameters. These masks are then injected into a pretrained video generation model via a ControlNet-style pathway, which aligns the action control signals with predicted videos, adds view-specific conditioning to accommodate camera viewpoints, and yields a unified world model architecture across embodiments. To mitigate overfitting to static backgrounds, BridgeV2W further introduces a flow-based motion loss that focuses on learning dynamic and task-relevant regions. Experiments on single-arm (DROID) and dual-arm (AgiBot-G1) datasets, covering diverse and challenging conditions with unseen viewpoints and scenes, show that BridgeV2W improves video generation quality compared to prior state-of-the-art methods. We further demonstrate the potential of BridgeV2W on downstream real-world tasks, including policy evaluation and goal-conditioned planning. More results can be found on our project website at https://BridgeV2W.github.io .

具身智能视频生成控制网络机器人建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。