用多智能体评审机制提升GUI智能体的奖励准确性与可扩展性
OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards
- 将任务轨迹分解为可验证节点,逐级审核决策证据链
- 在线强化学习中提升10.3%,自训练中提升6.9%性能
- 适用于需要高可靠性奖励的跨平台GUI智能体开发
强化学习有望提升GUI智能体在随机环境中的鲁棒性,但训练对奖励函数质量极为敏感。现有奖励方法难以兼顾可扩展性与性能。为此,我们提出OS-Themis——一个可扩展且准确的多智能体评判框架。不同于单一评判者,OS-Themis将轨迹分解为可验证里程碑,隔离关键决策证据,并通过审查机制严格审计证据链后作出最终判断。为便于评估,我们进一步构建OmniGUIRewardBench(OGRBench),一个面向跨平台GUI结果奖励的综合性基准。所有测试模型在OS-Themis下均达到最佳表现。在AndroidWorld上的大量实验表明,使用OS-Themis进行在线强化学习训练时性能提升10.3%,在自训练循环中用于轨迹验证与筛选时提升6.9%,充分展现了其推动智能体演进的潜力。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet training is highly sensitive to the quality of the reward function. Existing reward approaches struggle to achieve both scalability and performance. To address this, we propose OS-Themis, a scalable and accurate multi-agent critic framework. Unlike a single judge, OS-Themis decomposes trajectories into verifiable milestones to isolate critical evidence for decision making and employs a review mechanism to strictly audit the evidence chain before making the final verdict. To facilitate evaluation, we further introduce OmniGUIRewardBench (OGRBench), a holistic cross-platform benchmark for GUI outcome rewards, where all evaluated models achieve their best performance under OS-Themis. Extensive experiments on AndroidWorld show that OS-Themis yields a 10.3% improvement when used to support online RL training, and a 6.9% gain when used for trajectory validation and filtering in the self-training loop, highlighting its potential to drive agent evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。