arXiv:2510.06780cs.CLcs.AI2025-10Conference of the …被引 3

研究大模型知识提取的终止性与可靠性,发现结果基本可复现但受模型和语言影响。

Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness

  • 用小规模子数据集测试知识提取是否能停止、能否重复、是否稳定。
  • 提取过程大多能终止,但不同模型效果差异大。
  • 对种子和温度变化敏感,对语言和模型切换较脆弱,适合关注知识质量的研究者。

大型语言模型(LLMs)蕴含大量事实知识,但如何测量并系统化这些知识仍具挑战。将知识转化为结构化形式,例如通过递归提取方法如GPTKB(Hu等,2025b),尚处于探索阶段。关键开放问题包括:此类提取能否终止?输出是否可复现?对扰动是否稳健?本文使用miniGPTKB(领域特定、可处理的小型子数据集),在历史、娱乐和金融三个领域中,系统分析了终止性、可复现性和鲁棒性,涵盖产出量、词汇相似度和语义相似度三类指标。实验对比了四种变量(种子、语言、随机性、模型),结果显示:(i) 终止率高,但依赖于模型;(ii) 可复现性表现不一;(iii) 稳健性因扰动类型而异:对种子和温度变化高,对语言和模型切换较低。这表明大模型知识提取能可靠揭示核心知识,但也暴露出重要局限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) encode substantial factual knowledge, yet measuring and systematizing this knowledge remains challenging. Converting it into structured format, for example through recursive extraction approaches such as the GPTKB methodology (Hu et al., 2025b), is still underexplored. Key open questions include whether such extraction can terminate, whether its outputs are reproducible, and how robust they are to variations. We systematically study LLM knowledge materialization using miniGPTKBs (domain-specific, tractable subcrawls), analyzing termination, reproducibility, and robustness across three categories of metrics: yield, lexical similarity, and semantic similarity. We experiment with four variations (seed, language, randomness, model) and three illustrative domains (from history, entertainment, and finance). Our findings show (i) high termination rates, though model-dependent; (ii) mixed reproducibility; and (iii) robustness that varies by perturbation type: high for seeds and temperature, lower for languages and models. These results suggest that LLM knowledge materialization can reliably surface core knowledge, while also revealing important limitations.

大模型知识提取可复现性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。