用形式化方法生成高质量网页信息获取数据,提升智能体表现
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- 通过集合论形式化信息获取任务,用知识投影精确控制推理结构
- 多步扩展生成复杂问题,合成数据在GAIA和WebWalkerQA上达顶尖水平
- 适合研究智能体数据合成与开放域问答的学者参考
大型语言模型驱动的智能体通过基于网络的信息获取能力,推动了人工智能在复杂开放任务上的发展。然而,高质量训练数据的缺乏限制了信息获取智能体的进展。现有方法通常采用信息驱动范式,先收集网页数据再生成问题,导致信息结构与推理结构不一致、问答匹配度低。为此,我们提出形式化驱动的信息获取数据合成框架WebShaper,通过集合论系统化形式化信息获取任务。核心是知识投影(KP)概念,支持通过KP操作组合精准控制推理结构。合成过程从种子任务出发,经多步扩展:每一步由智能体扩增器利用检索与验证工具,基于形式化规则生成更复杂的正式问题。我们在合成数据上训练模型,实验表明,WebShaper在GAIA和WebWalkerQA基准上达到开源智能体的最先进性能。
原文摘要 · Abstract (English)
The advent of Large Language Model (LLM)-powered agents has revolutionized artificial intelligence by enabling solutions to complex, open-ended tasks through web-based information-seeking (IS) capabilities. The scarcity of high-quality training data has limited the development of IS agents. Existing approaches typically adopt an information-driven paradigm that first collects web data and then generates questions based on the retrieval. However, this may lead to inconsistency between information structure and reasoning structure, question and answer. To mitigate, we propose a formalization-driven IS data synthesis framework WebShaper to construct a dataset. WebShaper systematically formalizes IS tasks through set theory. Central to the formalization is the concept of Knowledge Projections (KP), which enables precise control over reasoning structure by KP operation compositions. During synthesis, we begin by creating seed tasks, then use a multi-step expansion process. At each step, an agentic Expander expands the current formal question more complex with retrieval and validation tools based on our formalization. We train our model on the synthesized dataset. Experiment results demonstrate that WebShaper achieves state-of-the-art performance among open-sourced IS agents on GAIA and WebWalkerQA benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。