提升GUI导航的跨域泛化能力,通过结构化推理与历史摘要增强智能体表现。
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- 引入结构化推理链,结合进展评估与决策分析生成连贯思考过程。
- 在标准基准上达到顶尖性能,跨域场景下效果显著优于现有方法。
- 适合研究多模态大模型、GUI自动化及强化学习应用的开发者和研究人员。
尽管多模态大语言模型(MLLMs)推动了GUI导航智能体的发展,但当前方法在跨领域泛化和有效利用历史信息方面仍存在局限。我们提出一种增强推理的框架,系统整合结构化推理、动作预测与历史摘要。结构化推理模块生成融合进展估计与决策推理的连贯思维链,既指导即时动作预测,又生成紧凑的历史摘要以支持后续步骤。基于此框架,我们通过监督微调伪标签轨迹并结合组相对策略优化(GRPO)进行强化学习,训练出名为GUI-Rise的GUI智能体。该框架采用专用奖励机制,包括关注历史的客观目标,直接将摘要质量与后续动作表现挂钩。在标准基准上的全面评估显示,在相同训练数据条件下达到最先进水平,尤其在跨域场景中表现突出。结果验证了该框架在多样化GUI导航任务中保持鲁棒推理与泛化的能力。代码已开源:https://leon022.github.io/GUI-Rise。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, \textbf{GUI-Rise}, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Code is available at https://leon022.github.io/GUI-Rise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。