arXiv:2505.13909cs.AIcs.CL2025-05被引 15

用少量人工轨迹+AI生成,让电脑操作智能体训练更高效

Efficient Agent Training for Computer Use

  • 仅需312条人工轨迹,结合AI合成多样化动作决策
  • 在新基准上相对提升141%,超越原模型10%
  • 适合想低成本训练高质量电脑操作智能体的研究者

大规模高质量轨迹数据长期是构建类人电脑操作智能体的关键瓶颈。我们提出PC Agent-E框架,显著降低对大规模人类示范的依赖。仅用312条人工标注的电脑操作轨迹,通过Claude 3.7 Sonnet合成多样替代动作决策进行数据增强。在该增强数据上训练的PC Agent-E模型,在我们新发布的WindowsAgentArena-V2基准上实现了141%的相对性能提升,甚至在相对指标上超越Claude 3.7 Sonnet 10%。该方法融合了可靠的人类电脑操作技能与自动化AI数据合成能力,不仅大幅优于仅使用人类轨迹训练的效果,也明显优于直接从Claude 3.7 Sonnet进行知识蒸馏的结果。代码、数据和模型已开源。

原文摘要 · Abstract (English)

Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesizing diverse alternative action decisions with Claude 3.7 Sonnet. Trained on these enriched trajectories, our PC Agent-E model achieved a remarkable 141 relative improvement, and even surpassed the Claude 3.7 Sonnet by 10% in relative terms on WindowsAgentArena-V2, an improved benchmark we also released. By integrating robust human computer use skills with automated AI data synthesis capabilities, our method not only brought substantial improvements over training on human trajectories alone, but also significantly surpassed direct distillation from Claude 3.7 Sonnet. Code, data and models are available at https://github.com/GAIR-NLP/PC-Agent-E

智能体训练数据合成电脑操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。