arXiv:2607.05468cs.ROcs.AI2026-07被引 5

让机器人模型学会看动作中的4维几何变化,提升操作精度且不增加计算负担。

Learning 4D Geometric Priors for Inference-Efficient World Action Models

论文配图:Learning 4D Geometric Priors for Inference-Efficient World Action Models
图 1 · 摘自论文原文
  • 用4维几何先验引导视频与动作联合训练,保留轻量推理结构。
  • 在LIBERO和真实场景中表现优异,最高达98.2%成功率达,推理成本不变。
  • 适合需要高精度动作规划的机器人系统,尤其擅长复杂物理交互任务。

世界动作模型(WAMs)通过联合建模视觉未来动态与可执行动作序列,在机器人操作中展现出强大潜力。然而,现有视频-动作联合训练方法主要优化外观导向的视频隐变量,难以充分捕捉精确操作所需的时序几何变化。我们提出MECo-WAM,一种多专家协同训练的世界动作模型,将与动作相关的4D几何先验注入视频-动作表征,同时保持原有轻量级推理结构。训练阶段,MECo-WAM结合视频与动作专家,并引入一个由冻结的VGGT编码器提供关系目标的轻量4D专家;通过非对称专家可见性防止辅助几何信息对动作生成产生非因果捷径。为将几何知识迁移到部署路径中,我们设计衰减式4D读掩码注意力机制,在训练初期提供有限帧几何引导,逐步消除依赖。此外,提出动作感知的时间几何蒸馏方法,对齐帧内几何关系及其时序演化,强化与机器人动作最相关的视觉区域。部署时,所有辅助4D组件均被移除。在LIBERO(98.2%)、RoboTwin 2.0(92.6%)及挑战性真实世界操作任务上的实验表明,MECo-WAM在不增加推理开销的前提下显著提升操作性能。

原文摘要 · Abstract (English)

World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.

机器人操作4D几何动作模型轻量化推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。