提出双系统框架,让界面智能体像人一样边做边学。
Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents
- 用快速视觉解析+相对奖励优化双系统,模拟人类决策过程。
- 在新基准上性能超越现有方法,多任务泛化能力更强。
- 适合研究人机交互、自动化工具的开发者参考。
图形用户界面(GUI)智能体通过计算机视觉和语言模型在自动化数字任务上取得显著进展。然而,现有系统存在明显局限:主要依赖试错式决策,缺乏从交互中学习与适应的能力;同时评估指标过于简单,难以反映真实界面交互的复杂性。本文提出CogniGUI认知框架,旨在克服这些挑战,实现类人化的自适应学习。受卡尼曼双系统理论启发,该方法结合两个核心组件:(1) 全能解析引擎,通过快速视觉语义分析对界面元素进行即时层次化解析,识别可操作项;(2) 基于群体相对策略优化(GRPO)的基底智能体,采用独特相对奖励机制评估多种操作路径,引导选择最简高效的行为序列。这种双系统设计支持迭代的“探索-学习-精炼”循环,使智能体能基于经验持续优化策略。此外,为评估系统的泛化与适应能力,我们引入ScreenSeek基准,涵盖多应用导航、动态状态转换及跨界面一致性等常被忽略的挑战。实验表明,CogniGUI在现有基准及新提出的基准上均优于当前最优方法。
原文摘要 · Abstract (English)
Graphical User Interface (GUI) agents have made significant progress in automating digital tasks through the utilization of computer vision and language models. Nevertheless, existing agent systems encounter notable limitations. Firstly, they predominantly depend on trial and error decision making rather than progressive reasoning, thereby lacking the capability to learn and adapt from interactive encounters. Secondly, these systems are assessed using overly simplistic single step accuracy metrics, which do not adequately reflect the intricate nature of real world GUI interactions. In this paper, we present CogniGUI, a cognitive framework developed to overcome these limitations by enabling adaptive learning for GUI automation resembling human-like behavior. Inspired by Kahneman's Dual Process Theory, our approach combines two main components: (1) an omni parser engine that conducts immediate hierarchical parsing of GUI elements through quick visual semantic analysis to identify actionable components, and (2) a Group based Relative Policy Optimization (GRPO) grounding agent that assesses multiple interaction paths using a unique relative reward system, promoting minimal and efficient operational routes. This dual-system design facilitates iterative ''exploration learning mastery'' cycles, enabling the agent to enhance its strategies over time based on accumulated experience. Moreover, to assess the generalization and adaptability of agent systems, we introduce ScreenSeek, a comprehensive benchmark that includes multi application navigation, dynamic state transitions, and cross interface coherence, which are often overlooked challenges in current benchmarks. Experimental results demonstrate that CogniGUI surpasses state-of-the-art methods in both the current GUI grounding benchmarks and our newly proposed benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。