arXiv:2508.04337cs.CLcs.AI2025-08综述被引 5

为文献综述的句子角色标注设计新标准,提升AI自动生成综述的能力。

Modelling and Classifying the Components of a Literature Review

  • 提出可机器处理的清晰标注框架,统一文献句子的修辞角色分类。
  • 在700句专家标注+2240句自动生成数据上,微调后模型F1超96%。
  • 小模型通过合成数据增强也能表现优异,适合资源有限的研究者使用。

以往研究表明,对科学论文中的句子按其修辞角色(如研究空白、结果、局限性、方法扩展等)进行标注,能显著提升AI分析科学文献的效果,并有望支持新一代高质量文献综述生成系统的发展。然而,实现这一目标需要定义合理的标注体系和高效的大规模标注策略。本文从两方面应对挑战:一是提出一种专为可靠自动处理设计的新颖、无歧义的标注框架;二是全面评估多种大语言模型(LLMs)在此框架下的修辞角色分类性能。为此,我们构建了Sci-Sentence基准,包含700句领域专家手动标注和2240句由LLM自动标注的句子。我们在该基准上评估了37个不同家族与规模的LLM,涵盖零样本学习与微调两种方法。实验表明,经过高质量数据微调后,现代大模型表现强劲,F1值超过96%,包括GPT-4o等大型专有模型以及轻量级开源模型均表现良好。此外,通过引入半合成的LLM生成示例扩充训练集,可进一步提升性能,使小型编码器获得稳健表现,并显著改进多个开源解码模型。

原文摘要 · Abstract (English)

Previous work has demonstrated that AI methods for analysing scientific literature benefit significantly from annotating sentences in papers according to their rhetorical roles, such as research gaps, results, limitations, extensions of existing methodologies, and others. Such representations also have the potential to support the development of a new generation of systems capable of producing high-quality literature reviews. However, achieving this goal requires the definition of a relevant annotation schema and effective strategies for large-scale annotation of the literature. This paper addresses these challenges in two ways: 1) it introduces a novel, unambiguous annotation schema that is explicitly designed for reliable automatic processing, and 2) it presents a comprehensive evaluation of a wide range of large language models (LLMs) on the task of classifying rhetorical roles according to this schema. To this end, we also present Sci-Sentence, a novel multidisciplinary benchmark comprising 700 sentences manually annotated by domain experts and 2,240 sentences automatically labelled using LLMs. We evaluate 37 LLMs on this benchmark, spanning diverse model families and sizes, using both zero-shot learning and fine-tuning approaches. The experiments reveal that modern LLMs achieve strong results on this task when fine-tuned on high-quality data, surpassing 96% F1, with both large proprietary models such as GPT-4o and lightweight open-source alternatives performing well. Moreover, augmenting the training set with semi-synthetic LLM-generated examples further boosts performance, enabling small encoders to achieve robust results and substantially improving several open decoder models.

文献分析修辞标注大模型评测知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。