arXiv:2604.15671cs.RO2026-04

让机器人在实验室长期任务中积累经验,提升复杂实验成功率。

Long-Term Memory for VLA-based Agents in Open-World Task Execution

  • 双层记忆架构保存成功操作路径,实现经验复用。
  • 在复杂化学实验中任务成功率显著高于现有模型。
  • 适合需要长期记忆与自主优化的智能机器人场景。

视觉-语言-动作(VLA)模型在具身决策方面展现出巨大潜力,但在复杂化学实验室自动化中的应用仍受限于长时程推理能力不足和缺乏持续经验积累。现有框架通常将规划与执行分离,难以固化成功策略,导致多阶段协议中试错效率低下。本文提出ChemBot,一种双层闭环框架,结合自主AI代理与进度感知的VLA模型(Skill-VLA),实现层级任务分解与执行。ChemBot采用双层记忆架构,将成功轨迹转化为可检索资产;通过模型上下文协议(MCP)服务器实现子代理与工具的高效调度。为克服VLA模型固有缺陷,进一步引入基于未来状态的异步推理机制,缓解轨迹不连续问题。在协作机器人上的大量实验表明,ChemBot在复杂长周期化学实验中,相较现有VLA基线,在操作安全性、精度和任务成功率方面均表现更优。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated significant potential for embodied decision-making; however, their application in complex chemical laboratory automation remains restricted by limited long-horizon reasoning and the absence of persistent experience accumulation. Existing frameworks typically treat planning and execution as decoupled processes, often failing to consolidate successful strategies, which results in inefficient trial-and-error in multi-stage protocols. In this paper, we propose ChemBot, a dual-layer, closed-loop framework that integrates an autonomous AI agent with a progress-aware VLA model (Skill-VLA) for hierarchical task decomposition and execution. ChemBot utilizes a dual-layer memory architecture to consolidate successful trajectories into retrievable assets, while a Model Context Protocol (MCP) server facilitates efficient sub-agent and tool orchestration. To address the inherent limitations of VLA models, we further implement a future-state-based asynchronous inference mechanism to mitigate trajectory discontinuities. Extensive experiments on collaborative robots demonstrate that ChemBot achieves superior operational safety, precision, and task success rates compared to existing VLA baselines in complex, long-horizon chemical experimentation.

机器人长时记忆VLA实验自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。