提出自迭代框架Co-EPG,让GUI智能体的规划与定位能力共同进化。
Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
- 通过循环反馈机制,规划与定位模型交替优化彼此。
- 仅三轮迭代即超越现有方法,无需外部数据。
- 适合研究GUI自动化与自进化AI系统的学者。
图形用户界面(GUI)任务自动化是人工智能研究的关键前沿。尽管高效的GUI智能体需协同整合规划与定位能力,但当前方法存在两大局限:(1) 未能充分挖掘跨模型协同效应;(2) 过度依赖合成数据生成而利用率不足。为此,我们提出Co-EPG——一种用于规划与定位协同进化的自迭代训练框架。Co-EPG建立了一个迭代正向反馈环:在该环中,规划模型通过基于定位奖励引导的组相对策略优化(GRPO)探索更优策略,生成多样化数据以优化定位模型;同时,优化后的定位模型为规划模型后续的GRPO训练提供更有效奖励,推动持续改进。因此,Co-EPG通过自对弈优化与训练数据提炼实现智能体能力的迭代提升。在Multimodal-Mind2Web和AndroidControl基准上,我们的框架在仅三轮迭代后即超越现有最先进方法,且无需外部数据。智能体在每轮迭代中均持续进步,展现出强大的自我增强能力。本工作确立了GUI智能体的新训练范式,从孤立优化转向集成、自驱动的协同进化方式。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current methodologies exhibit two fundamental limitations: (1) insufficient exploitation of cross-model synergies, and (2) over-reliance on synthetic data generation without sufficient utilization. To address these challenges, we propose Co-EPG, a self-iterative training framework for Co-Evolution of Planning and Grounding. Co-EPG establishes an iterative positive feedback loop: through this loop, the planning model explores superior strategies under grounding-based reward guidance via Group Relative Policy Optimization (GRPO), generating diverse data to optimize the grounding model. Concurrently, the optimized Grounding model provides more effective rewards for subsequent GRPO training of the planning model, fostering continuous improvement. Co-EPG thus enables iterative enhancement of agent capabilities through self-play optimization and training data distillation. On the Multimodal-Mind2Web and AndroidControl benchmarks, our framework outperforms existing state-of-the-art methods after just three iterations without requiring external data. The agent consistently improves with each iteration, demonstrating robust self-enhancement capabilities. This work establishes a novel training paradigm for GUI agents, shifting from isolated optimization to an integrated, self-driven co-evolution approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。