提出分离引用意图与内容类型的框架,提升论文引用分类准确率。
Semantically Orthogonal Framework for Citation Classification: Disentangling Intent and Content
- 将引用意图与被引内容类型解耦,基于语义角色理论设计新标注体系。
- 在跨领域数据集上,模型与人工标注一致性提升,分类性能优于现有框架。
- 适合数字图书馆、学术分析系统,推动科研评估标准化。
理解引用的作用对科研评估和引用感知的数字图书馆至关重要。然而,现有引用分类框架常将引用意图(为何引用)与被引内容类型(引用了什么)混淆,导致自动分类效果受限,难以兼顾细粒度区分与实际可靠性。本文提出SOFT框架,通过两个维度显式分离引用意图与被引内容类型,借鉴语义角色理论。我们系统性地重新标注ACL-ARC数据集,并发布一个从ACT2中采样的跨学科测试集。使用零样本和微调的大语言模型进行评估,结果表明SOFT显著提升了人类标注者与大模型之间的一致性,支持更强的分类性能和跨领域泛化能力,优于ACL-ARC和SciCite标注框架。这些结果证实SOFT作为清晰可复用的标注标准具有重要价值,有助于提升数字图书馆与学术传播基础设施的清晰性、一致性和通用性。所有代码与数据均已开源至GitHub:https://github.com/zhiyintan/SOFT。
原文摘要 · Abstract (English)
Understanding the role of citations is essential for research assessment and citation-aware digital libraries. However, existing citation classification frameworks often conflate citation intent (why a work is cited) with cited content type (what part is cited), limiting their effectiveness in auto classification due to a dilemma between fine-grained type distinctions and practical classification reliability. We introduce SOFT, a Semantically Orthogonal Framework with Two dimensions that explicitly separates citation intent from cited content type, drawing inspiration from semantic role theory. We systematically re-annotate the ACL-ARC dataset using SOFT and release a cross-disciplinary test set sampled from ACT2. Evaluation with both zero-shot and fine-tuned Large Language Models demonstrates that SOFT enables higher agreement between human annotators and LLMs, and supports stronger classification performance and robust cross-domain generalization compared to ACL-ARC and SciCite annotation frameworks. These results confirm SOFT's value as a clear, reusable annotation standard, improving clarity, consistency, and generalizability for digital libraries and scholarly communication infrastructures. All code and data are publicly available on GitHub https://github.com/zhiyintan/SOFT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。