arXiv:2409.06433cs.DLcs.CL2024-09被引 4

用大模型和知识图谱,让学术文章贡献更易被组织和理解。

Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization

  • 用大模型结合认知知识图谱,结构化整理论文贡献。
  • 通过微调和提示工程,提升论文分类与关系推荐准确率。
  • 适合政策制定者、产业界人士快速获取跨领域学术信息。

每年发表的学术论文超过250万篇,研究人员难以跟上科学进展。将论文贡献整合到一种新型认知知识图谱(CKG)中,可超越标题和摘要提供的信息,有效组织学术知识。本研究利用大语言模型(LLMs)对学术论文进行分类,并以结构化、可比方式描述其贡献。尽管以往研究局限于特定领域,但大模型具备的广泛领域无关知识,为生成结构化贡献描述提供了巨大潜力。同时,通过提示工程或微调,可灵活使用小型高效、低成本且环保的大模型。我们的方法融合了大模型知识与经过领域专家验证的CKG数据,显著提升在论文分类和谓词推荐等任务上的表现。通过一种新颖的提示技术注入CKG知识,进一步提高学术知识提取精度。该方法已集成至开放研究知识图谱(ORKG),实现对学术知识的精准访问,有力支持跨领域的学术交流与传播,惠及政策制定者、工业从业者及公众。

原文摘要 · Abstract (English)

The increasing amount of published scholarly articles, exceeding 2.5 million yearly, raises the challenge for researchers in following scientific progress. Integrating the contributions from scholarly articles into a novel type of cognitive knowledge graph (CKG) will be a crucial element for accessing and organizing scholarly knowledge, surpassing the insights provided by titles and abstracts. This research focuses on effectively conveying structured scholarly knowledge by utilizing large language models (LLMs) to categorize scholarly articles and describe their contributions in a structured and comparable manner. While previous studies explored language models within specific research domains, the extensive domain-independent knowledge captured by LLMs offers a substantial opportunity for generating structured contribution descriptions as CKGs. Additionally, LLMs offer customizable pathways through prompt engineering or fine-tuning, thus facilitating to leveraging of smaller LLMs known for their efficiency, cost-effectiveness, and environmental considerations. Our methodology involves harnessing LLM knowledge, and complementing it with domain expert-verified scholarly data sourced from a CKG. This strategic fusion significantly enhances LLM performance, especially in tasks like scholarly article categorization and predicate recommendation. Our method involves fine-tuning LLMs with CKG knowledge and additionally injecting knowledge from a CKG with a novel prompting technique significantly increasing the accuracy of scholarly knowledge extraction. We integrated our approach in the Open Research Knowledge Graph (ORKG), thus enabling precise access to organized scholarly knowledge, crucially benefiting domain-independent scholarly knowledge exchange and dissemination among policymakers, industrial practitioners, and the general public.

知识图谱大模型学术组织提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。