针对论文引言中引用文献的学术定位问题,提出全新评估基准。
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
- 从引用决策角度构建评估框架,而非仅看文本相似度。
- 基于4万篇论文和109万文档,测试集含100篇论文相关工作段落。
- 揭示现有模型在引用选择与逻辑组织上的系统性缺陷,适合研究AI写作的学者。
大型语言模型在科学写作中表现出强语言流畅性,但相关工作生成(RWG)的评估仍有限。现有评估多沿用摘要类指标,以参考文献段落的词汇或语义相似性作为质量代理。然而,相关工作写作本质上是引用层面的学术定位任务:需选择、组织并框定前人研究,以阐明目标论文与已有研究的关系、差异及贡献。因此,模型可能生成连贯且语义相关的内容,却存在引用不当或位置错误等学术性失误,传统指标无法捕捉。为此,我们提出 extbf{RWGBench},一个从引用决策视角评估RWG的基准。该基准基于40,108篇计算机科学论文与109万条检索文档构建,包含经精心筛选的100篇论文及其已发表的相关工作段落。我们设计多维度评估框架,涵盖引用选择、上下文恰当性、结构组织与话语结构。实验揭示当前系统在标准评估下被掩盖的系统性缺陷,而Oracle研究进一步区分了检索与生成环节的瓶颈。人工评估显示,我们的引用中心指标与专家判断高度一致,优于表面文本指标。RWGBench为开发与评估更符合学术写作实践的引言生成系统提供了引用中心的测试平台。
原文摘要 · Abstract (English)
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit summarization-oriented metrics, using lexical or semantic similarity to reference sections as proxies for quality. However, related work writing is fundamentally a citation-level scholarly positioning task: it requires selecting, organizing, and framing prior work to clarify how a target paper relates to, differs from, and contributes beyond existing research.As a result, models may generate coherent and semantically-relevant text while exhibiting academically critical failures, such as inappropriate citation selection or misplaced references, that conventional metrics do not capture.To this end, we introduce \textbf{RWGBench}, a benchmark that evaluates RWG from the perspective of citation decision-making rather than text similarity. RWGBench is constructed from a large-scale collection of 40,108 computer science papers and a retrieval corpus of 1.09 million documents, with a carefully curated test set comprising 100 papers and their corresponding published related work sections.We propose a multi-dimensional evaluation framework that assesses citation selection, contextual appropriateness, organization, and discourse structure.Experiments reveal systematic limitations in current systems that are obscured by standard evaluations, while Oracle studies further disentangle retrieval-level and generation-level bottlenecks. Human evaluation further shows that our citation-centric metrics align substantially better with expert judgment than surface-level text metrics. RWGBench offers a citation-centric testbed for developing and evaluating related work generation systems that are better aligned with scholarly writing practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。