arXiv:2607.19343cs.CVcs.RO2026-07被引 1

用像素级动作掩码让视频模型同时实现正向预测与逆向推理。

Masked Visual Actions for Unified World Modeling

论文配图:Masked Visual Actions for Unified World Modeling
图 1 · 摘自论文原文
  • 用部分可见的物体运动轨迹作为控制信号,直接对接视频模型的视觉空间。
  • 仅用15小时真实与仿真数据微调,即可在多种场景和机械臂上保持高精度。
  • 支持机器人规划、策略评估和逆向动作生成,适合具身智能研究者。

视频模型蕴含丰富的视觉世界运动、交互与接触响应先验,是机器人世界建模的理想基础。核心挑战在于如何以与模型所学视觉空间对齐且仍具物理意义的方式传达动作。本文提出「掩码视觉动作」,一种以像素空间表达的动作接口,将动作表示为视频中任意实体的部分可见运动轨迹。揭示机器人运动时,模型作为前向动力学模型预测场景对低层机器人动作的响应;揭示期望物体运动时,同一模型可恢复与该结果一致的机器人行为。仅需15小时真实视频与仿真中的掩码样本进行微调,单个模型即可在多样场景和多种机器人形态下实现强视觉保真度与可控性。在下游操作任务中,模型生成的想象轨迹结果与真实执行高度相关,可用于策略评估;通过排序候选未来提升基于模型的决策能力;并支持逆向建模,从期望物体运动合成机器人动作。

原文摘要 · Abstract (English)

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.

世界建模视频生成机器人控制动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。