arXiv:2602.20200cs.ROcs.AI2026-02中稿 · CVPR被引 16

用双记忆机制提升机器人操作的效率与鲁棒性

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

  • 引入全局先验记忆和局部一致性记忆,优化动作生成
  • 在三个仿真环境上成功率最高达98.6%,真实世界长任务提升52.4%
  • 适合追求高效、稳定机器人控制的开发者与研究者

层级化视觉-语言-动作(VLA)模型已成为机器人操作的主流范式,通常由视觉-语言主干网络感知理解,结合生成式策略生成动作。然而其性能日益受限于动作生成过程:(i)推理效率低,因各向同性噪声先验与目标动作分布存在显著差距,导致去噪步骤增多且不可行样本频发;(ii)鲁棒性差,现有策略仅依赖当前观测,忽视历史序列约束,缺乏对任务进展与时间一致性的感知。为此,我们提出OptimusVLA,一种具有全局先验记忆(GPM)和局部一致性记忆(LCM)的双记忆VLA框架。GPM用从语义相似轨迹中检索的任务级先验替代高斯噪声,缩短生成路径并减少函数评估次数(NFE)。LCM动态建模已执行动作序列以推断任务进展,并注入学习到的一致性约束,强化轨迹的时间连贯性与平滑性。在三个仿真基准测试中,OptimusVLA持续优于强基线:在LIBERO上平均成功率98.6%,在CALVIN上优于pi_0 13.5%,在RoboTwin 2.0 Hard上达38%平均成功率。真实世界评估中,其在泛化与长时序任务上均排名第一,分别超越pi_0 42.9%与52.4%,同时实现2.9倍推理速度提升。

原文摘要 · Abstract (English)

Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bottlenecked by the action generation proceess. (i) Low inference efficiency. A pronounced distributional gap between isotropic noise priors and target action distributions, which increases denoising steps and the incidence of infeasible samples. (ii) Poor robustness. Existing policies condition solely on the current observation, neglecting the constraint of history sequence and thus lacking awareness of task progress and temporal consistency. To address these issues, we introduce OptimusVLA, a dual-memory VLA framework with Global Prior Memory (GPM) and Local Consistency Memory (LCM). GPM replaces Gaussian noise with task-level priors retrieved from semantically similar trajectories, thereby shortening the generative path and reducing the umber of function evaluations (NFE). LCM dynamically models executed action sequence to infer task progress and injects a learned consistency constraint that enforces temporal coherence and smoothness of trajectory. Across three simulation benchmarks, OptimusVLA consistently outperforms strong baselines: it achieves 98.6% average success rate on LIBERO, improves over pi_0 by 13.5% on CALVIN, and attains 38% average success rate on RoboTwin 2.0 Hard. In Real-World evaluation, OptimusVLA ranks best on Generalization and Long-horizon suites, surpassing pi_0 by 42.9% and 52.4%, respectively, while delivering 2.9x inference speedup.

机器人操作多模态生成动作规划效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。