arXiv:2608.04765cs.ROcs.AI2026-08

用文字记忆提升视觉语言机器人长任务执行能力

Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models

论文配图:Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 引入显式语言记忆,将时序观察转为带逻辑的文本序列
  • 在仿真和真实机器人上成功完成复杂长程任务,成功率显著提升
  • 适合需要长期规划与可解释决策的智能机器人研究者

视觉-语言-动作(VLA)模型统一了视觉感知、语言理解和机器人控制。然而现有VLA模型在长时任务中仍面临诸多挑战:稀疏专家示范限制跨任务组合泛化;长时任务的非马尔可夫特性使仅依赖当前观测的策略难以保持时间一致性;有限的闭环纠错导致执行误差累积;端到端动作微调可能削弱视觉语言模型(VLM)骨干的高层语义表征。为此,我们提出一种分层长时VLA架构,包含显式语言记忆模块。核心思想是将离散时序观测转化为具有时序逻辑的连贯文本记忆序列。系统解耦为高层VLM与低层VLA:高层VLM通过视觉问答训练范式进行语义推理,低层VLA基于子任务指令与视觉观测执行精确连续控制。高层VLM利用前序记忆作为上下文锚点,递归更新语言记忆与子任务指令,实现长时执行中的持续时间追踪与动态修正。我们在多个仿真环境及真实机器人平台上验证该方法,结果表明显式语言记忆显著提升了VLA模型在复杂长时任务中的成功率与鲁棒性,并提供了可解释的决策过程语义描述。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations constrain cross-task compositional generalization; the non-Markovian nature of long-horizon tasks makes it difficult for policies conditioned only on current observations to maintain temporal consistency; limited closed-loop error correction allows execution errors to accumulate; and end-to-end action fine-tuning may weaken the high-level semantic representations of vision-language model (VLM) backbones. To address these issues, we propose a hierarchical long-horizon VLA architecture with an explicit language-memory module. The central idea is to convert discrete temporal observations into a coherent textual memory sequence with temporal logic. The system is decoupled into a high-level VLM and a low-level VLA: the high-level VLM performs semantic reasoning through a visual question answering training paradigm, while the low-level VLA executes precise continuous control conditioned on subtask instructions and visual observations. The high-level VLM recursively updates both language memory and subtask instructions using the previous memory as a contextual anchor, enabling persistent temporal tracking and dynamic correction during long-horizon execution. We evaluate the proposed method in multiple simulation environments and conduct sim-to-real experiments on a real robotic platform. The results demonstrate that explicit language memory improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.

视觉语言机器人长程规划语言记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。