让手机助手适应界面变化,靠记忆保持任务理解力
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
- 用双层记忆存储界面功能和任务意图,应对界面频繁更新
- 在线测试中性能超越基线模型,分布外场景仍稳定表现
- 适合需要长期运行的自动化手机操作场景
基于大模型的移动端GUI代理可实现自主任务执行,但界面外观频繁更新和工作流重构会导致依赖历史数据训练的代理失效。尽管界面呈现发生变化,其功能语义与任务意图本质上保持稳定。基于此洞察,我们提出MAGNET——一种以记忆驱动的自适应代理框架,包含双层记忆结构:静态记忆将多样视觉特征关联至稳定的功能语义,实现鲁棒的动作定位;过程记忆则捕捉跨不同工作流的稳定任务意图。我们设计动态记忆演化机制,通过优先更新高频访问知识持续优化两类记忆。在线基准AndroidWorld评估显示显著优于基线模型,离线测试也证实其在分布外场景下具有一致性提升。结果表明,利用界面变更中的稳定结构可有效提升代理在动态软件环境中的性能与泛化能力。
原文摘要 · Abstract (English)
Mobile GUI agents powered by large foundation models enable autonomous task execution, but frequent updates altering UI appearance and reorganizing workflows cause agents trained on historical data to fail. Despite surface changes, functional semantics and task intents remain fundamentally stable. Building on this insight, we introduce MAGNET, a memory-driven adaptive agent framework with dual-level memory: stationary memory linking diverse visual features to stable functional semantics for robust action grounding and procedural memory capturing stable task intents across varying workflows. We propose a dynamic memory evolution mechanism that continuously refines both memories by prioritizing frequently accessed knowledge. Online benchmark AndroidWorld evaluations show substantial improvements over baselines, while offline benchmarks confirm consistent gains under distribution shifts. These results validate that leveraging stable structures across interface changes improves agent performance and generalization in evolving software environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。