arXiv:2603.00503cs.CV2026-03被引 3

用双记忆机制提升长序列网页代理的决策效率与准确性。

M$^2$: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight Retrieval

  • 引入内部轨迹压缩与外部洞察检索双记忆,减少冗余信息。
  • 在多个数据集上提升成功率最高达19.6%,令牌消耗降低58.7%。
  • 无需训练,适合需要高效推理的复杂长任务场景。

基于多模态大语言模型的智能体在自主网页导航中展现出巨大潜力,但处理长时序任务仍是关键瓶颈。现有方法依赖大量数据收集与模型训练,仍面临计算成本高、推理能力不足的问题。为此,我们提出M²,一种无需训练的记忆增强框架,旨在提升上下文效率与决策鲁棒性。该方法采用双层记忆机制:内部记忆通过动态轨迹压缩将冗长交互历史简化为紧凑状态更新;外部记忆通过从离线洞察库中检索可操作指导来引导智能体。在WebVoyager和OnlineMind2Web上的广泛评估表明,M²持续优于基线,使Qwen3-VL-32B的准确率提升最高达19.6%,令牌消耗减少58.7%;而专有模型如Claude的准确率提升最高达12.5%,同时显著降低计算开销。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on extensive data collection and model training, yet still struggle with high computational costs and insufficient reasoning capabilities when facing complex, long-horizon scenarios. To address this, we propose M$^2$, a training-free, memory-augmented framework designed to optimize context efficiency and decision-making robustness. Our approach incorporates a dual-tier memory mechanism that synergizes Dynamic Trajectory Summarization (Internal Memory) to compress verbose interaction history into concise state updates, and Insight Retrieval Augmentation (External Memory) to guide the agent with actionable guidelines retrieved from an offline insight bank. Extensive evaluations across WebVoyager and OnlineMind2Web demonstrate that M$^2$ consistently surpasses baselines, yielding up to a 19.6% success rate increase and 58.7% token reduction for Qwen3-VL-32B, while proprietary models like Claude achieve accuracy gains up to 12.5% alongside significantly lower computational overhead.

网页代理双记忆轨迹压缩长时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。