小模型通过自主探索与回溯经验,显著提升网页任务规划能力。
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

- 用自主探索和回溯经验生成高质量训练数据
- 7B小模型在真实任务上达30.6%准确率,超大模型
- 适合资源有限但需跨网站泛化的网页自动化场景
多模态网络代理可协助人类完成重复性GUI操作,有效任务规划是将复杂任务分解为可执行动作的关键。相比商用大模型,小型开源多模态大模型(MLLM)成本低且更保护隐私,但存在规划能力弱和跨网站泛化能力差的问题。为此,我们提出规划经验探索与利用(PEEU)方法,通过自主探索环境发现经验,并利用回溯经验合成严格对齐的高层训练数据。为定量分析驱动性能的泛化行为,我们提出任务分解层次分析框架(TDHAF),系统研究三个任务粒度下的组合泛化能力:低、中、高。分析表明,掌握低层原子技能并不保证高层规划能力,而高层任务训练能带来更强的分布外(OOD)泛化。在真实世界基准测试中,我们的7B模型达到30.6%准确率,优于更大的Qwen2.5-VL-32B模型。结果表明,构建回溯高层任务并利用经验对小模型的分布外规划能力至关重要。
原文摘要 · Abstract (English)
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions. While small open source MLLMs are cost efficient and privacy preserving compared with commercial large models, they suffer from weak planning and limited cross website generalization. To address these limitations, we introduce the planning experience exploration and utilization (PEEU) method, which autonomously explores environments to discover experiences and utilizes hindsight experience to synthesize strictly aligned, high level training data. To quantitatively analyze the generalization behaviors driving this performance, we propose the task decomposition hierarchical analysis framework (TDHAF) to systematically study compositional generalization across three task granularities: low, middle and high levels. Our analysis reveals that mastering low level atomic skills does not guarantee high level planning competence, while high level task training yields stronger OOD generalization. Experiments on real world benchmarks demonstrate PEEU's superior effectiveness: our 7B model achieves 30.6% accuracy, outperforming the much larger Qwen2.5-VL-32B model. These demonstrate constructing hindsight high level tasks and leveraging experiences is crucial for OOD planning abilities of small MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。