arXiv:2511.12882cs.ROcs.AI2025-11被引 6

用多视角轨迹视频提升机器人动作预测一致性。

Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos

  • 用多视角轨迹视频替代低级动作指令,实现精准视觉运动预测。
  • 在双臂复杂场景中,动作执行精度和物体交互准确率显著提升。
  • 适合关注机器人具身智能与物理交互建模的研究者。

具身世界模型旨在通过视觉观测与动作预测和交互物理世界。然而,现有模型难以将低级动作(如关节位置)精确转化为预测帧中的机器人运动,导致与真实物理交互不一致。为此,我们提出MTV-World,一种引入多视角轨迹视频控制的具身世界模型。具体地,不直接使用低级动作,而是利用相机内参、外参及笛卡尔空间变换生成的轨迹视频作为控制信号。由于3D动作投影至2D图像会损失空间信息,单一视角难以准确建模交互,因此我们设计多视角框架以补偿信息损失,确保与真实物理世界高度一致。MTV-World基于多视角轨迹视频输入并以每视图初始帧为条件,预测未来帧。为进一步系统评估机器人运动精度与物体交互准确性,我们开发了基于多模态大模型与指代视频目标分割模型的自动评估流水线。为衡量空间一致性,将其建模为物体位置匹配问题,并采用杰卡德指数(Jaccard Index)作为评价指标。大量实验表明,MTV-World在复杂双臂场景中实现了精准控制执行与准确的物理交互建模。

原文摘要 · Abstract (English)

Embodied world models aim to predict and interact with the physical world through visual observations and actions. However, existing models struggle to accurately translate low-level actions (e.g., joint positions) into precise robotic movements in predicted frames, leading to inconsistencies with real-world physical interactions. To address these limitations, we propose MTV-World, an embodied world model that introduces Multi-view Trajectory-Video control for precise visuomotor prediction. Specifically, instead of directly using low-level actions for control, we employ trajectory videos obtained through camera intrinsic and extrinsic parameters and Cartesian-space transformation as control signals. However, projecting 3D raw actions onto 2D images inevitably causes a loss of spatial information, making a single view insufficient for accurate interaction modeling. To overcome this, we introduce a multi-view framework that compensates for spatial information loss and ensures high-consistency with physical world. MTV-World forecasts future frames based on multi-view trajectory videos as input and conditioning on an initial frame per view. Furthermore, to systematically evaluate both robotic motion precision and object interaction accuracy, we develop an auto-evaluation pipeline leveraging multimodal large models and referring video object segmentation models. To measure spatial consistency, we formulate it as an object location matching problem and adopt the Jaccard Index as the evaluation metric. Extensive experiments demonstrate that MTV-World achieves precise control execution and accurate physical interaction modeling in complex dual-arm scenarios.

具身智能视觉运动多视角建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。