对比多种大模型生成能力,揭示其在技能问题上的表现差异。
Characterising LLM-Generated Competency Questions: a Cross-Domain Empirical Study using Open and Closed Models
- 设计量化指标,跨领域分析大模型生成的技能问题
- 不同模型生成的问题在可读性与结构复杂度上差异显著
- 适合关注AI辅助知识工程的开发者与研究者
技能问题(CQs)是本体工程中需求获取的核心。它们以自然语言问题形式表达本体应满足的要求,传统上由本体工程师与领域专家通过人工协作完成。生成式AI可规模化自动化生成CQ,从而降低门槛、扩大参与方,提升本体工程的可及性。然而,由于大模型在参数量、任务专精和开放程度等方面差异巨大,需系统化地分析其生成的CQ在可读性、与输入文本的相关性以及结构复杂度等可观测属性。本文提出一套多维量化评估方法,基于多个明确用例与场景,对包括KimiK2-1T、LLama3.1-8B、LLama3.2-3B(开源)和Gemini 2.5 Pro、GPT 4.1(闭源)在内的多种模型进行实证分析。结果表明,模型性能呈现由使用场景决定的独特生成特征。
原文摘要 · Abstract (English)
Competency Questions (CQs) are a cornerstone of requirement elicitation in ontology engineering. CQs represent requirements as a set of natural language questions that an ontology should satisfy; they are traditionally modelled by ontology engineers together with domain experts as part of a human-centred, manual elicitation process. The use of Generative AI automates CQ creation at scale, therefore democratising the process of generation, widening stakeholder engagement, and ultimately broadening access to ontology engineering. However, given the large and heterogeneous landscape of LLMs, varying in dimensions such as parameter scale, task and domain specialisation, and accessibility, it is crucial to characterise and understand the intrinsic, observable properties of the CQs they produce (e.g., readability, structural complexity) through a systematic, cross-domain analysis. This paper introduces a set of quantitative measures for the systematic comparison of CQs across multiple dimensions. Using CQs generated from well defined use cases and scenarios, we identify their salient properties, including readability, relevance with respect to the input text and structural complexity of the generated questions. We conduct our experiments over a set of use cases and requirements using a range of LLMs, including both open (KimiK2-1T, LLama3.1-8B, LLama3.2-3B) and closed models (Gemini 2.5 Pro, GPT 4.1). Our analysis demonstrates that LLM performance reflects distinct generation profiles shaped by the use case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。