构建首个自动化能力问题生成评测基准,推动知识工程工具标准化评估
Bench4KE: Benchmarking Automated Competency Question Generation
- 基于API设计可扩展评测系统,统一评估生成能力问题的工具
- 涵盖17个真实项目数据集,用多种相似度指标量化生成质量
- 适合知识工程、LLM应用研究者,助力工具对比与复现
大型语言模型(LLMs)为知识工程(KE)自动化提供了新机遇。近年来已有多个基于LLM的自动能力问题(CQ)生成方法出现,但缺乏统一评估标准,影响研究可比性与可复现性。为此,本文提出Bench4KE——一个面向KE自动化的可扩展API基准系统,当前版本聚焦于自动CQ生成的评测。该系统包含来自17个真实知识工程项目的高质量黄金标准数据集,并采用多套相似度度量方法评估生成结果质量。我们对6个近期基于LLM的CQ生成系统进行了对比分析,建立了未来研究的基准线。Bench4KE还支持扩展至其他KE任务,如SPARQL查询生成、本体测试与起草。代码与数据集已开源,遵循Apache 2.0许可。
原文摘要 · Abstract (English)
The availability of Large Language Models (LLMs) presents a unique opportunity to reinvigorate research on Knowledge Engineering (KE) automation. This trend is already evident in recent efforts developing LLM-based methods and tools for the automatic generation of Competency Questions (CQs), natural language questions used by ontology engineers to define the functional requirements of an ontology. However, the evaluation of these tools lacks standardization. This undermines the methodological rigor and hinders the replication and comparison of results. To address this gap, we introduce Bench4KE, an extensible API-based benchmarking system for KE automation. The presented release focuses on evaluating tools that generate CQs automatically. Bench4KE provides a curated gold standard consisting of CQ datasets from 17 real-world ontology engineering projects and uses a suite of similarity metrics to assess the quality of the CQs generated. We present a comparative analysis of 6 recent CQ generation systems, which are based on LLMs, establishing a baseline for future research. Bench4KE is also designed to accommodate additional KE automation tasks, such as SPARQL query generation, ontology testing and drafting. Code and datasets are publicly available under the Apache 2.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。