对比三种提问方法,发现大模型可快速生成问题但需人工优化。
A Comparative Study of Competency Question Elicitation Methods from Ontology Requirements
- 用人工、模板和大模型三种方式生成知识库提问
- 大模型生成的问题可接受度高但需进一步修改
- 首次构建多标注者对比数据集,适合知识工程研究者
能力问题(CQs)在知识工程中至关重要,指导本体的设计、验证与测试。现有文献提出了从完全人工到基于大语言模型(LLM)的多种表述方法,但对其输出特征的系统性比较仍较少。本文对三种不同CQ生成方法进行实证对比:由本体工程师人工编写、使用CQ模式实例化、以及采用前沿大模型生成。基于文化遗产领域的需求文档,分别生成CQ,并从可接受性、模糊性、相关性、可读性和复杂性五个维度评估。研究贡献包括:(i) 首个基于同一来源、不同方法生成的多标注者CQ数据集;(ii) 系统性比较各方法生成结果的特性。结果显示,不同方法生成的CQ具有不同特征,大模型可作为初始提问获取手段,但其输出依赖具体模型,且通常需后续优化才能用于需求建模。
原文摘要 · Abstract (English)
Competency Questions (CQs) are pivotal in knowledge engineering, guiding the design, validation, and testing of ontologies. A number of diverse formulation approaches have been proposed in the literature, ranging from completely manual to Large Language Model (LLM) driven ones. However, attempts to characterise the outputs of these approaches and their systematic comparison are scarce. This paper presents an empirical comparative evaluation of three distinct CQ formulation approaches: manual formulation by ontology engineers, instantiation of CQ patterns, and generation using state of the art LLMs. We generate CQs using each approach from a set of requirements for cultural heritage, and assess them across different dimensions: degree of acceptability, ambiguity, relevance, readability and complexity. Our contribution is twofold: (i) the first multi-annotator dataset of CQs generated from the same source using different methods; and (ii) a systematic comparison of the characteristics of the CQs resulting from each approach. Our study shows that different CQ generation approaches have different characteristics and that LLMs can be used as a way to initially elicit CQs, however these are sensitive to the model used to generate CQs and they generally require a further refinement step before they can be used to model requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。