arXiv:2412.08258cs.DLcs.AI2024-12中稿 · Information Proces…被引 25

用大模型自动构建工程领域学术本体,效果媲美顶尖模型且更省资源。

Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field

  • 基于IEEE术语表构建标准数据集,评估大模型识别四类语义关系的能力。
  • 部分小模型经提示工程优化后,性能接近大模型,F1最高达0.967。
  • 适合需要低成本构建知识体系的研究者与系统开发者使用。

研究主题的本体对于结构化科学知识至关重要,有助于科研人员在海量文献中导航,并支撑搜索引擎与推荐系统等智能系统。然而,人工构建本体成本高、速度慢,且易过时或过于泛化。为此,本文对大语言模型(LLMs)识别不同研究主题间语义关系的能力进行了全面分析,这是构建本体的关键步骤。我们基于IEEE术语表构建了黄金标准数据集,用于评估模型在四种关系类型上的表现:上位、下位、等同与其它。研究评估了17种不同规模、开放/专有、全量/量化模型的性能,同时测试了四种零样本推理策略。结果表明,Mixtral-8x7B、Dolphin-Mistral-7B和Claude 3 Sonnet分别取得0.847、0.920和0.967的F1分数。此外,经过提示工程优化的小型量化模型,在显著降低计算开销的前提下,性能可媲美大型专有模型。

原文摘要 · Abstract (English)

Ontologies of research topics are crucial for structuring scientific knowledge, enabling scientists to navigate vast amounts of research, and forming the backbone of intelligent systems such as search engines and recommendation systems. However, manual creation of these ontologies is expensive, slow, and often results in outdated and overly general representations. As a solution, researchers have been investigating ways to automate or semi-automate the process of generating these ontologies. This paper offers a comprehensive analysis of the ability of large language models (LLMs) to identify semantic relationships between different research topics, which is a critical step in the development of such ontologies. To this end, we developed a gold standard based on the IEEE Thesaurus to evaluate the task of identifying four types of relationships between pairs of topics: broader, narrower, same-as, and other. Our study evaluates the performance of seventeen LLMs, which differ in scale, accessibility (open vs. proprietary), and model type (full vs. quantised), while also assessing four zero-shot reasoning strategies. Several models have achieved outstanding results, including Mixtral-8x7B, Dolphin-Mistral-7B, and Claude 3 Sonnet, with F1-scores of 0.847, 0.920, and 0.967, respectively. Furthermore, our findings demonstrate that smaller, quantised models, when optimised through prompt engineering, can deliver performance comparable to much larger proprietary models, while requiring significantly fewer computational resources.

本体生成大模型应用知识图谱工程领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。