arXiv:2410.24021cs.CL2024-10

用知识图谱嵌入识别文本间思想影响,效果优于传统方法。

Detecting text level intellectual influence with knowledge graph embeddings

  • 构建论文知识图谱,用图神经网络学习嵌入表示
  • 在判断文章间是否存在引用上,准确率显著提升
  • 可快速部署并针对特定领域微调,适合研究者使用

追踪思想传播与影响是历史学、文化分析、计算社会科学及科学学等多个领域的核心问题。本文收集开源期刊论文语料,利用Gemini大模型生成知识图谱表示,并基于已有方法与一种新型图神经网络嵌入模型,预测文章对之间是否存在引用关系。实验表明,该知识图谱嵌入方法在区分有引用与无引用的文章对方面表现更优;模型训练完成后运行高效,且可针对特定语料库进行微调,满足个性化研究需求。结果说明,知识图谱中概念间的关联类型蕴含揭示思想影响的潜在信息,未来通过分析文档级知识图谱以挖掘隐含结构,可能带来重要洞见。

原文摘要 · Abstract (English)

Introduction: Tracing the spread of ideas and the presence of influence is a question of special importance across a wide range of disciplines, ranging from intellectual history to cultural analytics, computational social science, and the science of science. Method: We collect a corpus of open source journal articles, generate Knowledge Graph representations using the Gemini LLM, and attempt to predict the existence of citations between sampled pairs of articles using previously published methods and a novel Graph Neural Network based embedding model. Results: We demonstrate that our knowledge graph embedding method is superior at distinguishing pairs of articles with and without citation. Once trained, it runs efficiently and can be fine-tuned on specific corpora to suit individual researcher needs. Conclusion(s): This experiment demonstrates that the relationships encoded in a knowledge graph, especially the types of concepts brought together by specific relations can encode information capable of revealing intellectual influence. This suggests that further work in analyzing document level knowledge graphs to understand latent structures could provide valuable insights.

知识图谱思想影响图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。