arXiv:2502.14192cs.CLcs.DL2025-02被引 5

用大模型少样本构建NLP学术知识图谱,关联论文与概念

NLP-AKG: Few-Shot Construction of NLP Academic Knowledge Graph Based on LLM

  • 基于大模型从论文中提取实体与关系,构建跨论文的语义网络
  • 从6万篇论文中提取62万实体和227万关系,覆盖NLP领域核心研究
  • 提出子图社区摘要方法,提升跨文献问答准确率,适合研究者参考

大语言模型已广泛用于科学论文的问答系统。为提升回答的专业性与准确性,许多研究采用外部知识增强。然而现有科学文献中的外部知识结构通常只关注论文实体或领域概念,忽略了论文间通过共享概念产生的内在联系,导致涉及论文与概念结合的问题时回答不够全面具体。为此,我们提出一种新知识图谱框架,通过论文内部语义元素和论文间引用关系,捕捉学术论文间的深层概念关联,构建关系网络。基于少样本知识图谱构建方法,利用大模型开发了面向自然语言处理领域的学术知识图谱NLP-AKG,从ACL Anthology中的60,826篇论文中提取620,353个实体和2,271,584条关系。在此基础上,提出‘子图社区摘要’方法,并在三个NLP科学文献问答数据集上验证其有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely applied in question answering over scientific research papers. To enhance the professionalism and accuracy of responses, many studies employ external knowledge augmentation. However, existing structures of external knowledge in scientific literature often focus solely on either paper entities or domain concepts, neglecting the intrinsic connections between papers through shared domain concepts. This results in less comprehensive and specific answers when addressing questions that combine papers and concepts. To address this, we propose a novel knowledge graph framework that captures deep conceptual relations between academic papers, constructing a relational network via intra-paper semantic elements and inter-paper citation relations. Using a few-shot knowledge graph construction method based on LLM, we develop NLP-AKG, an academic knowledge graph for the NLP domain, by extracting 620,353 entities and 2,271,584 relations from 60,826 papers in ACL Anthology. Based on this, we propose a 'sub-graph community summary' method and validate its effectiveness on three NLP scientific literature question answering datasets.

知识图谱大模型NLP少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。