arXiv:2510.13913cs.CLcs.AI2025-10被引 3

通过逐步提升任务难度生成高质量网页智能体训练数据

Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms

  • 设计渐进式难度增强机制生成复杂问答对
  • 数据量更小但工具调用多样性翻倍,性能更强
  • 适合研究长程推理与智能体训练的数据构建

基于网页的深度研究智能体需通过与在线工具的长期交互解决复杂问答任务。现有方法多依赖知识图谱构建指令微调数据集,但缺乏对难度和质量的精细控制,难以体现长程推理所需复杂性。此外,多数研究混淆了数据与训练策略的影响,难以评估数据本身的有效性。本文提出双路径数据合成流程:通过逐步增加任务复杂度直至基准网页智能体失败来生成问答对。该基准智能体在过程中承担尝试、事实验证、答案对比和过滤等多重角色。为评估效果,采用从强智能体中蒸馏的受控训练设置。在多个网页基准测试中,尽管数据规模更小,本方法生成的数据仍使智能体表现优于现有数据集。尤其,其工具调用行为多样性提升一倍,模型训练后表现出更强性能且避免重复调用。

原文摘要 · Abstract (English)

Web-based 'deep research' agents aim to solve complex question - answering tasks through long-horizon interactions with online tools. These tasks remain challenging, as the underlying language models are often not optimized for long-horizon reasoning and exploration. Prior work has proposed workflows for constructing instruction-tuning datasets, often leveraging knowledge graphs. However, such methods typically lack fine-grained control over difficulty and quality, yielding synthetic data that falls short of capturing the complexity required for long-horizon reasoning. Furthermore, many studies conflate data and training effects by comparing models trained under different optimization recipes, making it difficult to isolate and evaluate the effectiveness of the data itself. We introduce a two-pronged data synthesis pipeline that generates question - answer pairs by progressively increasing task complexity until a frontier baseline web agent fails. The baseline agent plays multiple roles in this process: attempting the questions, validating factuality, checking for alternative answers, and enforcing filtering. To evaluate the effectiveness of our synthesis methods, we adopt a controlled training setup based on distillation from strong web agents. Experiments across multiple web-based benchmarks show that our dataset - despite being smaller - enables the training of more effective web agents than existing datasets. In particular, our data exhibits twice the diversity in tool-use actions, allowing models trained on it to achieve stronger performance while avoiding repetitive tool-calling behaviors.

智能体训练数据合成长程推理网页智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。