arXiv:2509.17917cs.AI2025-09被引 4

通过分步反馈强化学习,提升GUI智能体的推理与执行可靠性。

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

  • 用可验证原则约束奖励信号,保证推理过程可解释。
  • 在虚拟机中自动生成结构化交互轨迹,提升数据效率22.2%以上。
  • 适合需要高可靠性和复杂任务适应性的GUI自动化场景。

近期的GUI智能体在视觉理解与动作预测上表现优异,但面临奖励信号不可靠和在线轨迹生成能力有限的问题。本文提出Orcust框架,融合基于原则的奖励建模(PCRM)与在线虚拟机驱动的轨迹构建(OVTC),以增强交互式GUI任务中的推理可靠性与数据效率。该框架利用环境可验证原则与大模型生成的规则,生成可解释的奖励信号,约束长链思维推理与基于规则的反馈。OVTC通过注入式虚拟机自主收集带有显式过程与结构目标的结构化GUI交互轨迹,支持分步奖励模型训练,有效捕捉人类偏好并遵守任务约束。在涵盖感知对齐、基础操作与端到端任务执行的标准GUI基准测试中,Orcust性能达到当前最优,相较基线模型(Qwen2.5-VL-7B)在ScreenSpot和ScreenSpot-Pro上分别提升22.2%和23.9%。结果表明,Orcust能显著提升GUI智能体的推理能力、适应性与可扩展性。

原文摘要 · Abstract (English)

Recent advances in GUI agents have achieved remarkable grounding and action-prediction performance, yet existing models struggle with unreliable reward signals and limited online trajectory generation. In this paper, we introduce Orcust, a framework that integrates Principle-Constrained Reward Modeling (PCRM) and Online VM-Grounded Trajectory Construction (OVTC) to enhance reasoning reliability and data efficiency in interactive GUI tasks. We leverages environment-verifiable and LLM-derived principle to enforce interpretable reward signals that constrain long chain-of-thought reasoning and rule-based feedback. OVTC spins up instrumented virtual machines to autonomously collect structured GUI interaction trajectories with explicit procedural and structural objectives, enabling the training of a stepwise reward model that robustly captures human preferences and adheres to task-specific constraints. Extensive experiments on standard GUI benchmarks covering perceptual grounding, foundational operations, and end-to-end task execution reveal that Orcust achieves state-of-the-art performance, improving by 22.2\% on ScreenSpot and 23.9\% on ScreenSpot-Pro over the base model (i.e. Qwen2.5-VL-7B). The results demonstrate Orcust's effectiveness in enhancing the reasoning, adaptability and scalability of GUI agents across various environments and task complexities.

GUI智能体强化学习虚拟机分步反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。