arXiv:2602.14093cs.AIcs.LG2026-02被引 10

用可验证奖励自动生成高效图形界面训练环境,提升智能体泛化能力。

GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training

  • 通过多模态代码模型重构真实应用为轻量网页环境。
  • 环境延迟降低10倍,每轮训练节省超2.8万美元成本。
  • 支持自生成环境的自我进化,适合做智能体后训练的研究者。

在交互式环境中进行后训练是提升智能体泛化与长时程规划能力的关键。然而,真实应用训练受限于高延迟、难复现以及依赖噪声视觉代理的不可验证奖励。为此,我们提出GUI-GENESIS,首个自动合成具备可验证奖励的高效图形界面训练环境的框架。该框架利用多模态代码模型将真实应用重建为轻量级网页环境,并配备代码原生奖励——可执行断言,提供确定性奖励信号,消除视觉估计噪声。大量实验表明,相比真实应用训练,GUI-GENESIS使环境延迟降低10倍,每轮训练成本节省超过28,000美元。值得注意的是,使用GUI-GENESIS训练的智能体在未见的真实任务上表现优于基础模型14.54%,甚至超越真实世界强化学习基线3.27%。此外,我们发现模型能合成自身尚无法解决的环境,揭示了自改进智能体的潜在路径。

原文摘要 · Abstract (English)

Post-training GUI agents in interactive environments is critical for developing generalization and long-horizon planning capabilities. However, training on real-world applications is hindered by high latency, poor reproducibility, and unverifiable rewards relying on noisy visual proxies. To address the limitations, we present GUI-GENESIS, the first framework to automatically synthesize efficient GUI training environments with verifiable rewards. GUI-GENESIS reconstructs real-world applications into lightweight web environments using multimodal code models and equips them with code-native rewards, executable assertions that provide deterministic reward signals and eliminate visual estimation noise. Extensive experiments show that GUI-GENESIS reduces environment latency by 10 times and costs by over $28,000 per epoch compared to training on real applications. Notably, agents trained with GUI-GENESIS outperform the base model by 14.54% and even real-world RL baselines by 3.27% on held-out real-world tasks. Finally, we observe that models can synthesize environments they cannot yet solve, highlighting a pathway for self-improving agents.

GUI训练自动化环境强化学习可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。