arXiv:2606.20092cs.CV2026-06被引 13

让机器人记住关键视觉线索,解决长时间操作中视线被挡的问题。

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

论文配图:EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 用稀疏视觉证据记忆,动态预测并存储重要画面。
  • 在17个仿真和4个真实任务中,成功率平均提升40%。
  • 适合需要长期记忆的机器人操作场景,如复杂家务任务。

长时程机器人操作中,记忆仍是关键瓶颈,传统视觉-语言-动作(VLA)策略在任务线索被遮挡或无法观测时易失效。现有记忆增强方法或存在严重信息瓶颈,或因双系统解耦导致高延迟,或依赖无选择性的缓冲区积累大量视觉冗余。为此,我们提出EventVLA,一个基于稀疏视觉证据记忆的端到端框架,包含两个核心组件:基础视觉锚点用于保留初始和短期上下文,以及动态关键帧证据记忆(KEM)模块。KEM直接从VLA的潜在嵌入中预测未来关键帧概率,自主捕获并存储稀疏、任务关键的视觉事件。该前瞻性机制使策略能动态评估当前观测的未来因果效用,在视觉信息消失前保存其关键证据。此外,我们提出RoboTwin-MeM诊断基准,专门用于评估具有交互式视觉证据的非马尔可夫操作任务。大量实验表明,在17个需记忆的仿真任务和4个真实世界双臂任务中,EventVLA相比最先进记忆增强VLA平均成功率提升40%。

原文摘要 · Abstract (English)

Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse, task-critical visual events. This foresight-driven mechanism empowers the policy to dynamically evaluate the future causal utility of current observations, preserving transient visual evidence before it becomes unobservable. Furthermore, we propose RoboTwin-MeM, a diagnostic benchmark specifically designed to evaluate non-Markovian manipulation tasks with interactive visual evidence. Extensive evaluations show that across 17 memory-requiring simulation tasks and 4 real-world bimanual tasks, EventVLA achieves an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.

机器人操作视觉记忆长时程决策VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。