让GUI智能体持续学习新界面,解决动态变化下的任务漂移问题。
Continual GUI Agents
- 引入基于锚点的奖励机制,稳定界面元素随分辨率变化的定位
- 在多场景下超越现有方法,在12个新界面测试中保持90%以上准确率
- 适合需要长期运行的自动化测试或人机交互系统开发者
随着数字环境持续变化,新界面数据不断出现,带来新的领域或分辨率,使得在静态环境训练的智能体性能下降。本文提出持续学习的GUI智能体新任务,要求智能体在分布偏移的界面中持续适应。我们发现现有方法因用户界面交互点和区域的动态变化而失去稳定定位能力。为此,提出GUI-Anchoring in Flux(GUI-AiF)强化学习微调框架,包含两个新奖励:流变锚点奖励(APR-iF)与流变锚区奖励(ARR-iF),引导智能体对齐动态变化的交互点与区域,缓解现有策略过度依赖固定坐标或元素尺度等静态线索的问题。大量实验表明,GUI-AiF显著优于当前最优基线方法。本工作首次建立面向GUI智能体的持续学习框架,揭示了强化微调在持续界面学习中的巨大潜力。
原文摘要 · Abstract (English)
As digital environments (data distribution) are in flux, with new GUI data arriving over time-introducing new domains or resolutions-agents trained on static environments deteriorate in performance. In this work, we introduce Continual GUI Agents, a new task that requires GUI agents to perform continual learning under shifted domains and resolutions. We find existing methods fail to maintain stable grounding as GUI distributions shift over time, due to the diversity of UI interaction points and regions in fluxing scenarios. To address this, we introduce GUI-Anchoring in Flux (GUI-AiF), a new reinforcement fine-tuning framework that stabilizes continual learning through two novel rewards: Anchoring Point Reward in Flux (APR-iF) and Anchoring Region Reward in Flux (ARR-iF). These rewards guide the agents to align with shifting interaction points and regions, mitigating the tendency of existing reward strategies to over-adapt to static grounding cues (e.g., fixed coordinates or element scales). Extensive experiments show GUI-AiF surpasses state-of-the-art baselines. Our work establishes the first continual learning framework for GUI agents, revealing the untapped potential of reinforcement fine-tuning for continual GUI Agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。