arXiv:2601.05163cs.CL2026-01被引 11

开源文档问答代理,用工具驱动方式提升长文档理解能力。

DocDancer: Towards Agentic Document-Grounded Information Seeking

  • 构建工具驱动的代理框架,显式建模文档探索与理解过程。
  • 在两个长文本基准上表现优异,优于现有方法。
  • 提出数据合成流程,解决高质量训练数据稀缺问题。

文档问答(DocQA)聚焦于基于给定文档回答问题,但现有文档问答代理缺乏有效的工具利用能力,且主要依赖闭源模型。本文提出 DocDancer,一个端到端训练的开源文档问答代理。我们将 DocQA 视为信息搜寻问题,设计了一种工具驱动的代理框架,显式建模文档探索与理解过程。为支持此类代理的端到端训练,我们提出一种‘探索-再合成’的数据生成流水线,以缓解高质量训练数据的匮乏。在合成数据上训练后,模型在两个长上下文文档理解基准 MMLongBench-Doc 与 DocBench 上均展现出有效性。进一步分析揭示了智能体工具设计与合成数据生成的关键洞见。

原文摘要 · Abstract (English)

Document Question Answering (DocQA) focuses on answering questions grounded in given documents, yet existing DocQA agents lack effective tool utilization and largely rely on closed-source models. In this work, we introduce DocDancer, an end-to-end trained open-source Doc agent. We formulate DocQA as an information-seeking problem and propose a tool-driven agent framework that explicitly models document exploration and comprehension. To enable end-to-end training of such agents, we introduce an Exploration-then-Synthesis data synthesis pipeline that addresses the scarcity of high-quality training data for DocQA. Training on the synthesized data, the trained models on two long-context document understanding benchmarks, MMLongBench-Doc and DocBench, show their effectiveness. Further analysis provides valuable insights for the agentic tool design and synthetic data.

文档问答智能体开源长文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。