提出预操作纠错模型,提升GUI自动化决策可靠性。
Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- 设计预操作批评机制,提前评估动作后果与正确性。
- 在移动与网页端测试中,纠错准确率显著优于现有模型。
- 适合需要高可靠性的自动化流程,如金融交易或系统配置。
近年来,多模态大语言模型(MLLMs)被广泛用于多模态推理任务,包括图形用户界面(GUI)自动化。与一般离线多模态任务不同,GUI自动化在在线交互环境中执行,需基于环境实时状态进行逐步决策。该任务对每一步决策的容错率极低,任何错误可能累积导致不可逆后果,如删除数据或支付款项。为解决此问题,我们引入一种预操作批评机制,在实际执行前提供有效反馈,通过推理动作的潜在结果与正确性来实现。具体而言,我们提出一种建议感知梯度相对策略优化(S-GRPO)策略,构建预操作批评模型GUI-Critic-R1,引入新型建议奖励以增强反馈可靠性。此外,我们开发了一种基于推理自举的数据收集管道,创建了GUI-Critic-Train与GUI-Critic-Test数据集,填补了现有GUI批评数据的空白。静态实验在跨移动端与网页端的GUI-Critic-Test上显示,GUI-Critic-R1在批评准确率方面显著优于当前主流MLLMs。动态评估在GUI自动化基准上进一步验证了模型的有效性与优越性,表现为成功率提升与操作效率改善。
原文摘要 · Abstract (English)
In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-step decision-making based on real-time status of the environment. This task has a lower tolerance for decision-making errors at each step, as any mistakes may cumulatively disrupt the process and potentially lead to irreversible outcomes like deletions or payments. To address these issues, we introduce a pre-operative critic mechanism that provides effective feedback prior to the actual execution, by reasoning about the potential outcome and correctness of actions. Specifically, we propose a Suggestion-aware Gradient Relative Policy Optimization (S-GRPO) strategy to construct our pre-operative critic model GUI-Critic-R1, incorporating a novel suggestion reward to enhance the reliability of the model's feedback. Furthermore, we develop a reasoning-bootstrapping based data collection pipeline to create a GUI-Critic-Train and a GUI-Critic-Test, filling existing gaps in GUI critic data. Static experiments on the GUI-Critic-Test across both mobile and web domains reveal that our GUI-Critic-R1 offers significant advantages in critic accuracy compared to current MLLMs. Dynamic evaluation on GUI automation benchmark further highlights the effectiveness and superiority of our model, as evidenced by improved success rates and operational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。