arXiv:2606.03188cs.RO2026-06被引 5

通过几何与语义监督,提升世界动作模型的表征能力。

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models

论文配图:GeoSem-WAM: Geometry- and Semantic-Aware World Action Models
图 1 · 摘自论文原文
  • 引入几何与语义辅助预测分支,增强潜在表示
  • 在复杂场景中提升动作预测准确率与鲁棒性
  • 无需生成未来视频,推理高效适合实际应用

近期的世界动作模型(WAMs)在具身决策中表现优异,但其优势是否来自推理时的未来想象,还是预测训练带来的表征学习仍不明确。现有研究指出主要优势源于稳健的潜在表征,而非测试时生成未来观测。然而,现有WAMs多依赖基于RGB的未来预测,难以捕捉复杂环境的结构与空间信息。为此,本文提出一种结构化世界建模框架,通过几何与语义监督增强潜在表征。除未来RGB预测外,模型新增几何与语义表示的辅助预测分支,使统一潜在空间能同时捕获场景动态、空间几何与语义上下文。关键在于,该方法在测试时避免显式未来回溯或视频生成,保持高效推理。大量实验表明,结构化世界监督显著提升了动作预测精度、场景理解能力及在挑战性具身场景下的鲁棒性,展现了构建可扩展、高效WAMs的巨大潜力。

原文摘要 · Abstract (English)

Recent World Action Models (WAMs) have demonstrated impressive capabilities in embodied decision-making. However, whether their effectiveness stems from explicit future imagination during inference or representation learning induced by predictive training remains an open question. Emerging evidence suggests the primary advantage lies in learning robust latent representations rather than generating future observations at test time. Nevertheless, existing WAMs mainly rely on RGB-based future prediction, which provides limited structural and spatial understanding of complex environments. To address this, we propose a structured world modeling framework that enhances latent representations through geometric and semantic supervision. Alongside future RGB prediction, our model introduces two auxiliary prediction branches for future geometry and semantic representations, enabling it to jointly capture scene dynamics, spatial geometry, and semantic context within a unified latent space. Crucially, our approach preserves efficient inference by avoiding explicit future rollout or video generation at test time. Extensive experiments show that incorporating structured world supervision consistently improves action prediction accuracy, scene understanding, and robustness under challenging embodied scenarios, highlighting its potential for advancing scalable and efficient WAMs.

世界模型具身智能几何感知语义监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。