首个跨学科论文综述生成评估基准,揭示不同工具在各领域表现差异。
SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation
- 构建10个学科的1000篇人工撰写综述数据集,结合领域评分与参考对齐双重评估
- 发现深度研究类代理在所有学科中表现最稳定,综述生成系统结构规划最优
- 指出现有系统普遍缺乏高质量引用,适合需跨学科综述的科研人员参考
自动论文综述生成(ASG)旨在通过检索、组织和整合学术论文生成全面的文献综述。尽管专用ASG框架和深度研究代理发展迅速,现有评估仍集中于计算机科学或依赖通用标准,难以判断当前系统是否满足各学科的综述规范。我们提出SurveyLens,首个面向学科的ASG评估基准。SurveyLens包含SurveyLens-1k——一个涵盖10个学科的1000篇人工撰写的综述数据集,以及融合学科感知评分标准与参考对齐的双视角评估框架。在11个前沿系统(包括普通大模型、ASG系统和深度研究代理)上评估发现:深度研究代理是唯一在所有10个学科中表现稳定的范式;ASG系统在结构规划上领先;所有范式在引用质量方面均表现薄弱,为特定学科工具选择及未来ASG设计提供了实践指导。
原文摘要 · Abstract (English)
Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialized ASG frameworks and Deep Research agents, existing evaluations largely center on Computer Science or rely on generic criteria, leaving it unclear whether current systems satisfy the survey standards of diverse disciplines. We introduce SurveyLens, the first discipline-aware ASG benchmark. SurveyLens comprises SurveyLens-1k, a curated dataset of 1,000 human-written surveys across 10 disciplines, and a dual-lens framework that combines discipline-aware rubric scoring with reference-based alignment to human-written surveys. Evaluating 11 state-of-the-art systems across vanilla LLMs, ASG systems, and Deep Research agents, we find that Deep Research agents are the only paradigm robust across all 10 disciplines, ASG systems lead on structural planning, and all paradigms remain weak on reference quality, providing practical guidance for discipline-specific tool selection and future ASG design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。