用少量真实操作生成大量高质量桌面交互数据,提升自动化代理泛化能力
ANCHOR: Branch-Point Data Generation for GUI Agents
- 从少量真实操作中识别关键分支点,生成上下文相关的任务变体
- 新生成轨迹在多个基准上使模型性能超越零样本基线和现有合成方法
- 适合需要高泛化能力的桌面自动化研究与开发人员
端到端桌面GUI代理需要大量高质量交互数据,但人工示范成本高,现有合成方法常因任务多样性不足或轨迹噪声、目标漂移而受限。本文提出轨迹扩展框架Anchor,从少量经验证的种子示范出发,识别对应有意义状态变化的分支点,基于当前GUI上下文生成新的、状态感知的任务变体。执行代理遵循建议指令生成新轨迹,验证器通过状态感知检查与轨迹一致性约束确保任务完成。为提升监督质量,进一步采用任务条件化的步级过滤去除无依据动作,并对分支后段进行去噪以维持意图连贯性。在标准桌面基准OSWorld和WindowsAgentArena上的实验表明,使用扩展语料微调的模型在性能上持续优于零样本代理和代表性合成基线,并具备跨应用与操作系统泛化能力。
原文摘要 · Abstract (English)
End-to-end GUI agents for real desktop environments require large amounts of high-quality interaction data, yet collecting human demonstrations is expensive and existing synthetic pipelines often suffer from limited task diversity or noisy, goal-drifting trajectories. We present a trajectory expansion framework Anchor that bootstraps scalable desktop supervision from a small set of verified seed demonstrations. Starting from each seed, we identify branch points that correspond to meaningful state changes and propose new, state-grounded task variants conditioned on the current GUI context. An executing agent then follows the proposed instructions to generate new trajectories, while a verifier enforces task completion via state-aware checks and trajectory-level consistency. To improve supervision quality, we further apply task-conditioned step-level filtering to remove ungrounded actions and denoise post-branch segments to maintain coherent intent. Experiments on standard desktop benchmarks, OSWorld and WindowsAgentArena, show that models fine-tuned on our expanded corpus achieve consistent improvements over zero-shot agents and representative synthesis baselines, and generalize across applications and operating systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。