arXiv:2608.31118cs.AI2026-08

大模型规模对知识图谱构建效果不一,架构和来源比参数量更重要。

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

  • 对比13个大模型,统一流程评估术语分类、分类体系发现等任务表现
  • 9B到27B参数提升精度,但更大模型未必更好,不同任务差异明显
  • 模型架构和训练数据来源影响大于参数数量,选型需综合考量

大语言模型(LLM)规模对知识图谱学习(OL)性能的影响尚未充分阐明。本文在统一检索增强生成流程下,对来自Qwen3.5和Qwen3.6系列的13个密集与专家混合(MoE)模型,以及部分专有GPT版本进行了控制实验。所有模型使用相同的嵌入模型、检索配置、提示模板、解码设置、数据集与评估指标,在四个生物医学与材料科学领域的本体上评估术语类型标注、分类体系发现及非分类关系抽取。在密集型Qwen3.5系列中,参数量增加主要提升精度而非召回率,最大增益出现在9B至27B之间。然而,规模效应并非单调或一致,27B密集模型在术语类型标注上显著优于更大规模的稀疏模型,而更大的MoE模型在分类体系发现上表现最佳。非分类关系抽取在各模型中仍困难,尤其在材料数据科学本体中。匹配的Qwen变体与专有GPT版本间的性能差异表明,模型架构与训练谱系的影响可能超过名义参数量。研究结果表明,仅凭模型规模无法有效指导知识图谱学习,需结合架构与数据源进行选择。

原文摘要 · Abstract (English)

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.

知识图谱大模型评估本体学习模型规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。