评测大模型处理知识图谱能力,自动判断谁更懂语义技术。
LLM-KG-Bench 3.0: A Compass for SemanticTechnology Capabilities in the Ocean of LLMs
- 构建可扩展任务框架,自动化评估大模型对知识图谱的处理能力。
- 覆盖RDF/SPARQL、Turtle/JSON-LD等任务,测试30+主流模型表现。
- 支持开源与闭源模型,生成模型卡片对比性能,无需人工逐条核对。
当前大语言模型(LLM)能辅助编程等多项任务,但能否有效支持知识图谱(KG)工作?哪款模型在语义网络与知识图谱工程领域表现最佳?是否可在不依赖大量人工核查的前提下作出判断?本文推出的LLM-KG-Bench 3.0框架旨在回答上述问题。该框架包含一套可扩展的任务集,支持对大模型输出的自动化评估,覆盖语义技术多个方面。本文介绍该框架3.0版本,并提供基于多个前沿大模型生成的提示、答案及评估数据集。相比初版,新版本在任务API、任务设计、vllm库支持等方面实现显著升级,涵盖超过30款现代开源与专有大模型。通过该数据集,可生成典型模型卡片,展示其在处理RDF和SPARQL任务中的能力,并对比其在Turtle与JSON-LD格式序列化任务上的表现。
原文摘要 · Abstract (English)
Current Large Language Models (LLMs) can assist developing program code beside many other things, but can they support working with Knowledge Graphs (KGs) as well? Which LLM is offering the best capabilities in the field of Semantic Web and Knowledge Graph Engineering (KGE)? Is this possible to determine without checking many answers manually? The LLM-KG-Bench framework in Version 3.0 is designed to answer these questions. It consists of an extensible set of tasks for automated evaluation of LLM answers and covers different aspects of working with semantic technologies. In this paper the LLM-KG-Bench framework is presented in Version 3 along with a dataset of prompts, answers and evaluations generated with it and several state-of-the-art LLMs. Significant enhancements have been made to the framework since its initial release, including an updated task API that offers greater flexibility in handling evaluation tasks, revised tasks, and extended support for various open models through the vllm library, among other improvements. A comprehensive dataset has been generated using more than 30 contemporary open and proprietary LLMs, enabling the creation of exemplary model cards that demonstrate the models' capabilities in working with RDF and SPARQL, as well as comparing their performance on Turtle and JSON-LD RDF serialization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。