arXiv:2604.03181cs.ROcs.CV2026-04被引 7

用多视角3D热力图视频提升机器人操作的数据效率。

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

  • 设计多视角3D热力图与RGB视频联合预测模型,融合空间结构信息。
  • 仅需10条演示轨迹即完成训练,在3个基准上性能提升15%以上。
  • 适合追求少样本、强泛化能力的机器人控制研究者使用。

机器人操作需同时理解环境的三维空间结构和时间演化,但现有策略常忽视其一或两者。多数方法依赖二维视觉输入或基于静态图像-文本对预训练的骨干网络,导致数据需求高且对环境动态理解有限。为此,我们提出SpatialVAM,首个能同时预测空间感知多视角热力图视频与RGB视频的3D视频动作模型。核心思想是该设计自然将3D信息注入视频基础模型,并统一视频预训练与动作微调的表示格式。大量实验表明,SpatialVAM实现数据高效、鲁棒、可泛化且可解释的操作。仅需10条示范轨迹且无需额外预训练,即可处理复杂长时程与接触密集任务,泛化至分布外场景,并生成真实未来视频。在Meta-World(22%↑)、RoboCasa(15%↑)及真实机器人平台(16%↑)上的评估显示,SpatialVAM持续优于其他视频动作模型、视觉语言动作模型及3D基策略,建立数据高效多任务操作新基准。

原文摘要 · Abstract (English)

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. They often rely on 2D visual observations or backbones pretrained on static image--text pairs, which leads to high data requirements and limited comprehension of environment dynamics. To address this, we introduce SpatialVAM, the first 3D Video Action Model that simultaneously predict spatial-aware multi-view heatmap videos and RGB videos. Our key insight is that this design naturally injects 3D information into video foundation models while aligning the representation format between video pretraining and action finetuning. Extensive experiments demonstrate that SpatialVAM enables data-efficient, robust, generalizable, and interpretable manipulation. With only ten demonstration trajectories and no additional pretraining, SpatialVAM handles challenging long-horizon and contact-rich tasks, generalizes to out-of-distribution settings, and predicts realistic future videos. Evaluations on Meta-World (22\%$\uparrow$), RoboCasa (15\%$\uparrow$) and real-world robotic platforms (16\%$\uparrow$) show that SpatialVAM consistently outperforms other video action models, vision language action models and 3D-based policies, establishing a new state-of-the-art in data-efficient multi-task manipulation.

机器人控制视频生成多视角少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。