arXiv:2601.03540cs.CL2026-01ACL综述

评测大模型整合海量文献信息的能力,发现规划写作流程更靠谱

DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing

  • 用高质量综述论文反推研究需求,构建无检索干扰的合成测试集
  • 96项任务中,规划式生成比单次输出减少幻觉,结构更完整
  • 提出可量化的检查清单评估方法,让长文本合成有标准可依

大型语言模型向自主智能体演进推动了深度研究的发展。尽管检索能力已有良好基准,但后检索合成阶段——即模型需消化大量上下文并把零散证据整合为连贯长篇报告——因开放式写作的主观性而长期缺乏评估。为此,我们提出 DeepSynth-Eval,一个用于客观评估信息整合能力的基准。我们以高质量综述论文为金标准,逆向生成研究请求,并从其参考文献中构建“上帝视角上下文”,以剥离检索噪声。提出细粒度评估协议:使用通用检查清单(评估事实覆盖)与约束检查清单(评估结构组织),将主观判断转化为可验证指标。在96个任务上的实验表明,从数百份参考文献中整合信息仍是重大挑战。结果证明,采用规划-写作工作流的代理显著优于单次生成,有效降低幻觉并更好遵守复杂结构约束。

原文摘要 · Abstract (English)

The evolution of Large Language Models (LLMs) towards autonomous agents has catalyzed progress in Deep Research. While retrieval capabilities are well-benchmarked, the post-retrieval synthesis stage--where agents must digest massive amounts of context and consolidate fragmented evidence into coherent, long-form reports--remains under-evaluated due to the subjectivity of open-ended writing. To bridge this gap, we introduce DeepSynth-Eval, a benchmark designed to objectively evaluate information consolidation capabilities. We leverage high-quality survey papers as gold standards, reverse-engineering research requests and constructing "Oracle Contexts" from their bibliographies to isolate synthesis from retrieval noise. We propose a fine-grained evaluation protocol using General Checklists (for factual coverage) and Constraint Checklists (for structural organization), transforming subjective judgment into verifiable metrics. Experiments across 96 tasks reveal that synthesizing information from hundreds of references remains a significant challenge. Our results demonstrate that agentic plan-and-write workflows significantly outperform single-turn generation, effectively reducing hallucinations and improving adherence to complex structural constraints.

信息整合模型评估长文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。