用GAN模型精准识别引用意图,提升学术影响力分析精度
Leveraging GANs for citation intent classification and its impact on citation network analysis
- 基于GAN与上下文嵌入的引用意图分类方法
- 性能媲美顶尖模型,参数量大幅减少
- 发现引用类型过滤会显著改变论文网络地位
引用在科学生态系统中具有基础性作用,是知识传播、文献承认和学术影响力评估的核心。然而,并非所有引用功能相同:有的用于背景介绍,有的引入方法或比较结果。理解引用意图有助于更精细地解读学术影响力。本文采用基于GAN的方法进行引用意图分类,结果显示该方法性能接近当前最优水平,且参数量显著减少,验证了将GAN架构与上下文嵌入结合在意图分类任务中的有效性和高效性。我们进一步研究了过滤引用意图对引文网络中心性的影响。基于unArXiv数据集构建的网络分析表明,论文排名受引用意图影响显著。四种中心性指标——度数、PageRank、接近度和介数——均对引用类型过滤敏感,其中介数中心性表现最敏感,特定引用意图移除后排名发生显著变化。
原文摘要 · Abstract (English)
Citations play a fundamental role in the scientific ecosystem, serving as a foundation for tracking the flow of knowledge, acknowledging prior work, and assessing scholarly influence. In scientometrics, they are also central to the construction of quantitative indicators. Not all citations, however, serve the same function: some provide background, others introduce methods, or compare results. Therefore, understanding citation intent allows for a more nuanced interpretation of scientific impact. In this paper, we adopted a GAN-based method to classify citation intents. Our results revealed that the proposed method achieves competitive classification performance, closely matching state-of-the-art results with substantially fewer parameters. This demonstrates the effectiveness and efficiency of leveraging GAN architectures combined with contextual embeddings in intent classification task. We also investigated whether filtering citation intents affects the centrality of papers in citation networks. Analyzing the network constructed from the unArXiv dataset, we found that paper rankings can be significantly influenced by citation intent. All four centrality metrics examined- degree, PageRank, closeness, and betweenness - were sensitive to the filtering of citation types. The betweenness centrality displayed the greatest sensitivity, showing substantial changes in ranking when specific citation intents were removed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。