arXiv:2608.05706cs.CV2026-08

从多视角人类视频中学习3D感知的潜在动作,提升机器人世界模型泛化能力。

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

论文配图:LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
图 1 · 摘自论文原文
  • 通过多视角不变的动作分词与几何对齐约束,实现3D感知的潜在动作学习。
  • 在多个数据集上生成质量、物理一致性和泛化能力均达当前最优水平。
  • 适合研究机器人世界模型、自监督视觉表征与3D动作建模的开发者。

世界模型使智能体可在无需真实交互的情况下进行前向推演与规划。然而,其在开放世界具身智能中的应用受限于动作标注成本高及平台间动作空间异质性。近期,潜在动作模型(LAMs)通过自监督方式直接从无标注人类视频中学习动作表示,缓解了这一瓶颈。但多数现有LAMs依赖单视角输入,仅在2D像素空间操作,引发核心问题:仅将多视角视频引入训练能否赋予潜在动作3D感知?本研究表明答案是否定的,主因在于未来帧外观泄露、跨摄像机外观差异及视角变化。为此,我们提出LAWM-3D,包含三项紧密耦合的设计:(1) 多视角不变的统一动作分词方案,以学习3D感知的潜在动作;(2) 几何对齐约束,将中间编码特征锚定至预训练3D基础模型,显式提供跨视角几何对应关系;(3) 非单射的RGB-D联合重建目标,防止利用未来帧外观信息的捷径学习,迫使模型聚焦具有几何意义的运动线索。这些组件并非简单堆叠,而是基于统一动机紧密耦合。基于大规模人类视频预训练与机器人微调的两阶段范式,大量实验证明所提3D感知潜在动作显著提升世界模型性能,在生成质量、物理一致性与泛化能力方面均达到当前最优。

原文摘要 · Abstract (English)

World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recently, latent action models (LAMs) have alleviated this bottleneck by learning action representations directly from unlabeled human videos in a self-supervised manner. Nevertheless, most existing LAMs rely on single-view inputs and operate primarily in 2D pixel space, raising a fundamental question: can simply incorporating multi-view videos into LAM training endow the learned latent actions with 3D-aware perception? Our study shows that the answer is negative. The primary reasons lie in future-frame appearance leakage as well as inter-camera appearance discrepancies and viewpoint variations. To address these issues, we propose LAWM-3D, which introduces three tightly coupled key designs: (1) a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions; (2) a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, thereby explicitly providing cross-view geometric correspondences; and (3) a non-injective RGB-D joint reconstruction objective that prevents shortcut learning from future-frame appearance information, forcing the LAM to focus supervision on motion cues with geometric significance. Importantly, these components are not simply stacked but are tightly coupled through a unified motivation. Built upon a two-stage paradigm of large-scale human video pretraining followed by robot fine-tuning, extensive experiments demonstrate that the proposed 3D-aware latent actions significantly improve world model performance, achieving SOTA results in generation quality, physical consistency, and generalization ability.

世界模型3D感知动作学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。