用类脑结构记忆提升图形界面智能体的长流程任务能力
Hybrid Self-evolving Structured Memory for GUI Agents
- 构建图结构记忆库,融合符号节点与轨迹嵌入
- 支持多跳检索与运行时刷新,任务成功率提升22.5%
- 适合需要复杂交互的自动化工具开发人员
视觉语言模型(VLMs)的发展使图形界面智能体能以类人方式操作计算机。然而,现实任务仍因长周期流程、多样界面和频繁中间错误而困难。现有方法依赖外部记忆库,但仅采用扁平化检索,缺乏人类记忆的结构化组织与自进化特性。受大脑启发,我们提出混合自进化结构记忆(HyMEM),一种基于图的记忆系统,结合离散高阶符号节点与连续轨迹嵌入。该系统维持图结构,支持多跳检索、通过节点更新实现自进化,并在推理过程中实时刷新工作内存。大量实验表明,HyMEM持续提升开源GUI智能体性能,使7B/8B模型达到甚至超越强闭源模型水平;显著提升Qwen2.5-VL-7B达+22.5%,优于Gemini2.5-Pro-Vision和GPT-4o。
原文摘要 · Abstract (English)
The remarkable progress of vision-language models (VLMs) has enabled GUI agents to interact with computers in a human-like manner. Yet real-world computer-use tasks remain difficult due to long-horizon workflows, diverse interfaces, and frequent intermediate errors. Prior work equips agents with external memory built from large collections of trajectories, but relies on flat retrieval over discrete summaries or continuous embeddings, falling short of the structured organization and self-evolving characteristics of human memory. Inspired by the brain, we propose Hybrid Self-evolving Structured Memory (HyMEM), a graph-based memory that couples discrete high-level symbolic nodes with continuous trajectory embeddings. HyMEM maintains a graph structure to support multi-hop retrieval, self-evolution via node update operations, and on-the-fly working-memory refreshing during inference. Extensive experiments show that HyMEM consistently improves open-source GUI agents, enabling 7B/8B backbones to match or surpass strong closed-source models; notably, it boosts Qwen2.5-VL-7B by +22.5% and outperforms Gemini2.5-Pro-Vision and GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。