arXiv:2607.08182cs.CVcs.AI2026-07

让视觉语言动作模型聚焦关键信息,提升动态环境下的决策能力

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

论文配图:LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
图 1 · 摘自论文原文
  • 通过动态优先级机制引导模型关注任务相关区域
  • 在多个基准上超越现有方法,提升显著的泛化性能
  • 适合需要精准感知与推理的机器人控制场景

视觉-语言-动作(VLA)模型旨在将多模态输入映射为机器人动作。然而,现有方法因对所有视觉标记一视同仁且依赖人工选定因素,在复杂动态场景中表现不佳,缺乏对关键证据的强调和对潜在因素的捕捉。为此,我们提出LEEVLA,一种在潜在环境演化中识别关键信息的VLA架构,能显式引导模型关注有意义区域,同时保持潜在世界表征的结构演化。我们引入漂移引导的动态优先级(DGDP),结合动态位置优先级(DPP)与语义漂移引导(SDG),在训练中指导模型关注方向。此外,提出结构化特征流生成(SFFG),通过原型到边缘(P2P)预测建模优先特征在潜在空间中的演化,并使用互邻域对比(MC)损失维持邻域拓扑一致性。DGDP与SFFG共同构成任务感知的“何处-如何”训练框架。大量实验表明,LEEVLA在多个VLA基准上持续优于先前方法,证实显式任务证据引导与结构化潜在推理对可扩展VLA至关重要。代码已开源。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors. To address this issue, we propose LEEVLA, a VLA architecture for seeing what matters in Latent Environment Evolution that explicitly guides the model toward informative regions while preserving the structured evolution of latent world representations. To identify salient and instruction-relevant regions, we introduce drift-guided dynamic prioritization (DGDP), which combines dynamic position prioritization (DPP) with semantic drift guidance (SDG) to guide the VLA agent where to attend during training. On top of this, we introduce structured feature flow generation (SFFG), which models how these prioritized features should evolve in latent space via prototype-to-periphery (P2P) prediction, and a mutual-neighborhood contrastive (MC) loss to maintain topological consistency among neighborhoods. Together, DGDP and SFFG form a task-aware "where-how" training framework. Extensive experiments on VLA benchmarks show that LEEVLA consistently outperforms prior methods, confirming that explicit task-evidence guidance and structured latent reasoning are both crucial for scalable VLA. Our code is available at https://github.com/LyuQi127/LEEVLA.

机器人控制多模态学习注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。