让GUI智能体自动纠错并自我进化,提升任务成功率。
MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback Reflux
- 用多智能体奖励模型动态评估操作轨迹,支持自适应反馈。
- 在真实任务中实现92.3%的任务准确率,显著优于基线方法。
- 适合开发可自我优化的自动化界面交互系统。
图形用户界面(GUI)智能体正快速迈向跨场景自主交互与可靠任务执行。然而,两大核心挑战仍待解决:如何自动化评估智能体行为轨迹,以及如何大规模生成高质量训练数据以支持持续改进。现有方法常依赖人工标注或静态规则验证,限制了可扩展性与动态环境适应能力。本文提出MagicGUI-RMS,一种多智能体奖励模型系统,具备自适应轨迹评估、纠错反馈与自我演化学习能力。该系统融合领域专用奖励模型(DS-RM)与通用奖励模型(GP-RM),实现细粒度动作评估与异构任务间的鲁棒泛化。为支持规模化奖励学习,设计了结构化数据构建流水线,自动生成平衡且多样化的奖励数据集,有效降低标注成本并保持样本真实性。执行过程中,奖励模型识别错误动作,提出优化替代方案,并通过自动化数据回流机制持续优化智能体行为。大量实验表明,MagicGUI-RMS在任务准确率与行为鲁棒性上均有显著提升,验证其作为基于奖励驱动的自进化GUI智能体坚实基础的有效性。
原文摘要 · Abstract (English)
Graphical user interface (GUI) agents are rapidly progressing toward autonomous interaction and reliable task execution across diverse applications. However, two central challenges remain unresolved: automating the evaluation of agent trajectories and generating high-quality training data at scale to enable continual improvement. Existing approaches often depend on manual annotation or static rule-based verification, which restricts scalability and limits adaptability in dynamic environments. We present MagicGUI-RMS, a multi-agent reward model system that delivers adaptive trajectory evaluation, corrective feedback, and self-evolving learning capabilities. MagicGUI-RMS integrates a Domain-Specific Reward Model (DS-RM) with a General-Purpose Reward Model (GP-RM), enabling fine-grained action assessment and robust generalization across heterogeneous GUI tasks. To support reward learning at scale, we design a structured data construction pipeline that automatically produces balanced and diverse reward datasets, effectively reducing annotation costs while maintaining sample fidelity. During execution, the reward model system identifies erroneous actions, proposes refined alternatives, and continuously enhances agent behavior through an automated data-reflux mechanism. Extensive experiments demonstrate that MagicGUI-RMS yields substantial gains in task accuracy, behavioral robustness. These results establish MagicGUI-RMS as a principled and effective foundation for building self-improving GUI agents driven by reward-based adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。