用大模型模拟界面操作,低成本生成海量训练数据。
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- 用大模型生成多样化的界面状态与操作路径。
- 在两个平台上的表现媲美甚至超越真实数据训练的智能体。
- 针对性扩展策略可大幅降低对强模型依赖,适合高效训练新智能体。
数字智能体需大量、多样的用户界面轨迹才能泛化至真实任务,但收集此类数据在人力标注、基础设施和工程上成本极高。为此,我们提出 extbf{UI-Simulator},一种可扩展的范式,通过生成结构化界面状态与转换,在大规模上合成训练轨迹。该范式融合了用于多样化界面状态的数字世界模拟器、实现连贯探索的引导式滚动过程,以及生成高质量多样轨迹的轨迹包装器。我们进一步提出 extbf{UI-Simulator-Grow},一种目标导向的扩展策略,通过优先处理高影响任务,高效合成信息丰富的轨迹变体。在 WebArena 与 AndroidWorld 上的实验表明,尽管使用较弱教师模型,UI-Simulator 的鲁棒性显著优于基于真实界面训练的开源智能体。此外,仅以 Llama-3-8B-Instruct 为基础模型,UI-Simulator-Grow 即可达到使用 Llama-3-70B-Instruct 模型的性能,凸显了有针对性的合成扩展范式的潜力,可连续高效提升数字智能体能力。
原文摘要 · Abstract (English)
Digital agents require diverse, large-scale UI trajectories to generalize across real-world tasks, yet collecting such data is prohibitively expensive in both human annotation, infra and engineering perspectives. To this end, we introduce $\textbf{UI-Simulator}$, a scalable paradigm that generates structured UI states and transitions to synthesize training trajectories at scale. Our paradigm integrates a digital world simulator for diverse UI states, a guided rollout process for coherent exploration, and a trajectory wrapper that produces high-quality and diverse trajectories for agent training. We further propose $\textbf{UI-Simulator-Grow}$, a targeted scaling strategy that enables more rapid and data-efficient scaling by prioritizing high-impact tasks and synthesizes informative trajectory variants. Experiments on WebArena and AndroidWorld show that UI-Simulator rivals or surpasses open-source agents trained on real UIs with significantly better robustness, despite using weaker teacher models. Moreover, UI-Simulator-Grow matches the performance of Llama-3-70B-Instruct using only Llama-3-8B-Instruct as the base model, highlighting the potential of targeted synthesis scaling paradigm to continuously and efficiently enhance the digital agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。