arXiv:2606.28712cs.ROcs.LG2026-06

融合定位与动作条件预测,实现长期一致性建模。

J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor Graphs

  • 用联合因子图统一表示位置、预测状态和持久特征点
  • 在部分观测下提升长时预测一致性,减少漂移
  • 适合需要精准规划的机器人系统开发

传统同时定位与地图构建(SLAM)仅估计度量位姿和几何地图,无法提供动作条件的预测状态。动作条件世界模型虽能学习紧凑的隐状态动态,却忽略全局度量一致性,且在开环推演中产生累积漂移。本文提出J-LAW(联合定位与动作条件世界建模),一种统一的因子图框架,将度量位姿变量、预测隐状态与持续存在的隐特征点关联起来。J-LAW将每张图像表示为紧凑的预测状态,并通过独立学习的映射将其与位姿或运动测量结合。其最大后验(MAP)因子图在时间上强制这些互补信息源的一致性。在PushT和WildGS数据集上的实验表明,J-LAW的因子图表示可提升长时隐状态一致性,并在部分观测下恢复更可靠的预测状态,为未来的集成定位与规划系统奠定基础。

原文摘要 · Abstract (English)

Classical simultaneous localization and mapping (SLAM) estimates metric poses and a geometric map but does not provide an action-conditioned predictive state. Action-conditioned world models learn compact latent dynamics but ignore global metric consistency and accumulate drift under open-loop rollout. We introduce J-LAW (Joint Localization and Action-Conditioned World Modeling), a unified factor-graph formulation that connects metric pose variables, predictive latent states, and persistent latent landmarks in this letter.J-LAW represents each image as a compact predictive state and combines it with pose or motion measurements through a separately learned mapping. Its maximum a posteriori (MAP) factor graph enforces consistency between these complementary sources of information over time. Experiments on PushT and WildGS show that J-LAW's factor-graph representation can improve long-horizon latent consistency and recover more reliable predictive states under partial observations, forming a foundation for future integrated localization and planning systems.

机器人世界模型定位建图因子图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。