用知识图谱的原型知识,揭示大模型如何记忆与泛化信息。
Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs
- 将知识图谱转化为三类原型知识:词汇、层级、拓扑。
- 通过知识激活任务验证,模型在问答中依赖特定原型知识。
- 适合研究模型记忆机制或提升知识推理能力的研究者。
我们提出‘原型知识’概念,用于形式化和度量大语言模型(LLMs)在预训练期间对编码知识图谱的词元序列的内化程度及其在推理时的利用方式。尽管大模型在预训练中能记忆大量词元序列,但其如何将这些记忆转化为可复用的知识并实现泛化仍是核心开放问题。我们据此将原型知识分为词汇、层级和拓扑三种类型,依据需激活的知识类型而定。通过知识激活任务(KATs)测量原型知识,分析其语义偏差等通用特性。进一步研究原型知识对Text-to-SPARQL性能的影响,通过不同提示策略测试输入条件下的表现。采用新分析框架,评估模型预测是否成功激活对应原型知识。该方法为探索语义级数据污染提供实用工具,并适用于闭源预训练模型的优化。
原文摘要 · Abstract (English)
We introduce the concept of protoknowledge to formalize and measure how sequences of tokens encoding Knowledge Graphs are internalized during pretraining and utilized at inference time by Large Language Models (LLMs). Indeed, LLMs have demonstrated the ability to memorize vast amounts of token sequences during pretraining, and a central open question is how they leverage this memorization as reusable knowledge through generalization. We then categorize protoknowledge into lexical, hierarchical, and topological forms, varying on the type of knowledge that needs to be activated. We measure protoknowledge through Knowledge Activation Tasks (KATs), analyzing its general properties such as semantic bias. We then investigate the impact of protoknowledge on Text-to-SPARQL performance by varying prompting strategies depending on input conditions. To this end, we adopt a novel analysis framework that assesses whether model predictions align with the successful activation of the relevant protoknowledge for each query. This methodology provides a practical tool to explore Semantic-Level Data Contamination and serves as an effective strategy for Closed-Pretraining models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。