arXiv:2607.11270cs.ROcs.AI2026-07被引 1

提出Lumo-2模型,通过结构化对齐实现可预测、可扩展的机器人学习。

Towards Predictive, Aligned, and Scalable Robot Learning

论文配图:Towards Predictive, Aligned, and Scalable Robot Learning
图 1 · 摘自论文原文
  • 在隐空间中推理世界动态生成动作,统一视觉与语言模态
  • 实测在长时序、精细操作等任务上显著优于基线模型
  • 适合研究具身智能、多模态控制及可扩展机器人系统的人

学习的本质在于推理和解决新问题的能力,而不仅限于记忆。我们提出Lumo-2,一种基于隐空间世界-动作建模的方法,通过在隐空间中推理世界动态来生成动作。所学隐空间世界动态捕捉物理相关的视觉变化,自然编码未来可能性,并为跨模态对齐提供统一基础。该方法实现类似世界建模的可预测推理,同时保持轻量且聚焦于控制相关的物理动态。核心假设是:动作生成质量由隐空间几何结构决定。我们发现,传统的重建导向动作标记化目标会诱导低层信号保真度偏差,导致重建质量与下游控制性能不一致。为此,提出分阶段的多模态预对齐策略,逐步将动作表示与隐空间世界动态、视觉和语言对齐。该过程强化跨模态一致性,促进抽象,构建利于预测推理的结构化隐空间。我们系统性地研究了隐空间世界建模与模态对齐的作用,分析其在缩放规律与分布外泛化中的影响。结果表明,Lumo-2在需要时间推理、物理理解或高控制复杂度的挑战性真实任务中持续超越强基线模型,包括长时程与灵巧操作任务。这些发现表明,结构化多模态对齐与可预测推理是推动具身智能发展的根本原则。

原文摘要 · Abstract (English)

Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to world modelling while remaining lightweight and focused on physical dynamics relevant to control. Central to our approach is the hypothesis that action generation quality is governed by the geometry of the latent space. We observe that standard reconstruction-based action tokenization objectives induce representations biased toward low-level signal fidelity, leading to misalignment between reconstruction quality and downstream control performance. To address this limitation, we propose a multi-stage modality pre-alignment strategy in which action representations are progressively aligned with latent world dynamics, vision, and language. This process enforces cross-modal consistency, promotes abstraction, and induces a structured latent space for predictive reasoning. We provide a systematic empirical study of latent world modelling and modality alignment, analyzing their roles in scaling laws and out-of-distribution generalization. Results show that Lumo-2 consistently outperforms strong vision-language-action (VLA) and world-action model (WAM) baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high control complexity, including long-horizon and dexterous manipulation. These findings suggest that structured multimodal alignment and predictive reasoning are fundamental principles for advancing embodied intelligence.

机器人学习多模态对齐隐空间建模具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。