用大模型自动验证知识图谱问题,省时省力还准确
Large Language Models Assisting Ontology Evaluation
- 用1393个问答对训练大模型,自动判断知识图谱是否满足需求
- o1-preview和o3-mini模型表现接近人工平均水准
- 可嵌入Protégé工具,辅助专家快速完成知识库验证
通过功能需求(如通过能力问题验证)评估知识图谱是成熟但成本高、耗时且易出错的过程,即使对专家也是如此。本文提出OE-Assist框架,实现能力问题的自动化与半自动化验证。基于包含1,393个能力问题及其对应知识图谱和背景故事的数据集,我们首次系统性地研究了大语言模型(LLM)在知识图谱评估中的应用,主要贡献包括:(i) 在人工标注的黄金标准上评估基于LLM的自动验证方法的有效性;(ii) 开发并评估一个基于LLM的框架,通过提供建议来辅助在Protégé中进行能力问题验证。实验表明,使用o1-preview和o3-mini模型的自动化评估性能与普通用户平均水平相当。
原文摘要 · Abstract (English)
Ontology evaluation through functional requirements, such as testing via competency question (CQ) verification, is a well-established yet costly, labour-intensive, and error-prone endeavour, even for ontology engineering experts. In this work, we introduce OE-Assist, a novel framework designed to assist ontology evaluation through automated and semi-automated CQ verification. By presenting and leveraging a dataset of 1,393 CQs paired with corresponding ontologies and ontology stories, our contributions present, to our knowledge, the first systematic investigation into large language model (LLM)-assisted ontology evaluation, and include: (i) evaluating the effectiveness of a LLM-based approach for automatically performing CQ verification against a manually created gold standard, and (ii) developing and assessing an LLM-powered framework to assist CQ verification with Protégé, by providing suggestions. We found that automated LLM-based evaluation with o1-preview and o3-mini perform at a similar level to the average user's performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。