arXiv:2509.26251cs.CV2025-09

让视觉语言动作模型看懂空间结构和动态变化,提升决策稳定性和可解释性。

Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models

  • 引入几何感知编码与多尺度时序建模,增强对空间结构和运动模式的理解。
  • 在仿真与真实场景中均达到领先性能,显著提升动作建模的鲁棒性。
  • 适合需要高可靠性与可解释性的机器人决策系统研究者参考。

潜在动作模型(LAMs)使视觉-语言-动作(VLA)系统能够从大规模无标注数据中学习语义动作表示。然而,我们发现两大瓶颈:1)常用的端到端训练图像编码器存在空间理解能力差的问题;2)当输入帧时间间隔较远时,LAMs表现脆弱,导致时序感知能力有限。为此,我们提出Farsighted-LAM,一种具有几何感知空间编码与多尺度时序建模的潜在动作框架,能从连续帧中捕捉结构先验与动态运动模式。进一步提出SSM-VLA,一个基于Farsighted-LAM的端到端VLA框架,融合结构化感知与视觉思维链模块,显式推理环境动态,增强决策一致性与可解释性。我们在多种仿真与真实世界任务中验证了SSM-VLA,取得当前最优表现。结果表明,结合几何感知、时序一致性与显式推理的策略,有效提升了具身智能的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Latent Action Models (LAMs) enable Vision- Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottlenecks of LAMs: 1) the commonly adopted end-to-end trained image encoder suffers from poor spatial understanding; 2) LAMs can be fragile when input frames are temporally distant, leading to limited temporal percep- tion. Such factors inevitably hinder stable and clear action modeling. To this end, we propose Farsighted-LAM, a latent action framework with geometry-aware spatial encoding and multi-scale temporal modeling, capturing structural priors and dynamic motion patterns from consecutive frames. We further propose SSM-VLA, an end-to-end VLA framework built upon Farsighted-LAM, which integrates structured perception with a visual Chain-of-Thought module to explicitly reason about environmental dynamics, enhancing decision consistency and interpretability. We validate SSM-VLA on multiple VLA tasks in both simulation and real-world settings, and achieve state-of- the-art performance. Our results demonstrate that our strategy of combining geometry-aware modeling, temporal coherence, and explicit reasoning is effective in enhancing the robustness and generalizability of embodied intelligence.

视觉语言动作具身智能时序建模几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。