arXiv:2510.03120cs.CL2025-10被引 4

用测验题评估大模型写综述是否符合读者需求

SurveyBench: Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?

  • 构建测验驱动的细粒度评测框架,以读者需求为导向
  • 基于1.1万篇arXiv论文和近5000份高质量综述训练数据
  • 发现现有大模型综述平均比人类差21%,评测结果具可解释性

学术综述写作需将海量文献提炼为连贯且有洞见的叙事,但该过程仍耗时费力。尽管近期出现通用DeepResearch代理和专用方法(统称LLM4Survey)可自动生成综述,其输出常不及人工水平,且缺乏严格、以读者为中心的评测基准来揭示其不足。为此,我们提出细粒度、测验驱动的评测框架SurveyBench,包含:(1) 来自最近11,343篇arXiv论文及对应4,947份高质量综述的典型综述主题;(2) 多层次指标体系,涵盖结构质量(如覆盖广度、逻辑连贯性)、内容质量(如综合粒度、洞察清晰度)以及非文本丰富性;(3) 双模式评测协议,包括基于内容和基于测验的答案可答性测试,明确对齐读者信息需求。结果显示,SurveyBench有效挑战现有LLM4Survey方法,在内容评测中平均比人类低21%。

原文摘要 · Abstract (English)

Academic survey writing, which distills vast literature into a coherent and insightful narrative, remains a labor-intensive and intellectually demanding task. While recent approaches, such as general DeepResearch agents and survey-specialized methods, can generate surveys automatically (a.k.a. LLM4Survey), their outputs often fall short of human standards and there lacks a rigorous, reader-aligned benchmark for thoroughly revealing their deficiencies. To fill the gap, we propose a fine-grained, quiz-driven evaluation framework SurveyBench, featuring (1) typical survey topics source from recent 11,343 arXiv papers and corresponding 4,947 high-quality surveys; (2) a multifaceted metric hierarchy that assesses the outline quality (e.g., coverage breadth, logical coherence), content quality (e.g., synthesis granularity, clarity of insights), and non-textual richness; and (3) a dual-mode evaluation protocol that includes content-based and quiz-based answerability tests, explicitly aligned with readers' informational needs. Results show SurveyBench effectively challenges existing LLM4Survey approaches (e.g., on average 21% lower than human in content-based evaluation).

大模型评测综述生成读者需求arXiv

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。