arXiv:2607.17738cs.CL2026-07

用大模型提升论文引用功能分类,效果显著优于以往方法。

Large Language Models for Citation Function Classification

论文配图:Large Language Models for Citation Function Classification
图 1 · 摘自论文原文
  • 对比五种大模型在零样本、少样本和微调下的表现
  • 微调后的Falcon 7B在ACL-ARC上达73.3%宏F1分数
  • 提出新数据集AC3,区分评价性引用与中性引用

引用功能分类在理解科学文献间关系与推动文献计量分析方面至关重要。本研究首次系统评估了多种前沿大语言模型(LLMs)在该任务上的表现,在ACL-ARC数据集上取得新的最优结果。我们对比了五种模型(Mistral 7B、Orca 2-7B、LLaMA 3.1-8B、Falcon 7B 和 SciBERT),涵盖零样本、少样本及微调三种策略。其中,微调后的Falcon 7B模型在ACL-ARC上达到73.3%的宏F1分数,显著优于此前方法。此外,我们构建了新数据集AC3,采用七类标注体系,能区分中性致谢与显性评价立场(如批评、赞扬、反驳等)。该数据集通过四种上下文提取变体实现,系统评估了上下文范围对分类性能的影响。研究还提供了模型表现、实验配置与局限性的详细分析,为该领域未来研究提供指导。据我们所知,这是首个专注于引用功能分类全面模型比较的研究,填补了近期综述中指出的空白。

原文摘要 · Abstract (English)

Citation function classification plays a crucial role in understanding the relationships between scientific publications and advancing bibliometric analysis. This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset. We systematically compare five models (Mistral 7B, Orca 2-7B, LLaMA 3.1-8B, Falcon 7B, and SciBERT) across zero-shot, few-shot, and fine-tuning approaches. Our fine-tuned Falcon 7B model achieves a 73.3% macro F1 score on ACL-ARC, representing a significant improvement over previous methods. Additionally, we introduce AC3, a novel dataset featuring a seven-category annotation scheme that distinguishes between neutral acknowledgments and explicit evaluative stances (more opinion-oriented citations - criticizing, complimenting, contradicting). The dataset is implemented across four context extraction variants to systematically evaluate the impact of contextual scope on classification performance. We also provide detailed analysis of model performance, experimental configurations, and limitations to guide future research in this domain. To our knowledge, this is one of the first studies dedicated to comprehensive model comparison for citation function classification, addressing a gap identified in recent surveys.

大模型引用分类文本分类自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。