从图结构视角分析大模型知识分布,发现邻近实体知识相似性。
A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models
- 将大模型知识建模为图,量化三元组与实体层面的知识
- 发现拓扑相近实体具有相似知识水平,存在知识同质性
- 基于邻居信息预测实体知识,可用于筛选低知三元组微调
大型语言模型被广泛研究作为神经知识库,在知识访问、可编辑性、推理和可解释性方面表现突出。然而,极少工作关注其知识的结构模式。为此,我们从图视角探究这些结构特征,量化了大模型在三元组和实体层面的知识,并分析其与节点度等图结构属性的关系。进一步发现知识同质性:拓扑上接近的实体具有相似的知识水平,这促使我们构建基于图机器学习的模型,通过局部邻居估计实体知识。该模型可用于识别大模型不熟悉的三元组,进行知识检查。实验证明,使用筛选出的三元组进行微调可带来更优性能。
原文摘要 · Abstract (English)
Large language models have been extensively studied as neural knowledge bases for their knowledge access, editability, reasoning, and explainability. However, few works focus on the structural patterns of their knowledge. Motivated by this gap, we investigate these structural patterns from a graph perspective. We quantify the knowledge of LLMs at both the triplet and entity levels, and analyze how it relates to graph structural properties such as node degree. Furthermore, we uncover the knowledge homophily, where topologically close entities exhibit similar levels of knowledgeability, which further motivates us to develop graph machine learning models to estimate entity knowledge based on its local neighbors. This model further enables valuable knowledge checking by selecting triplets less known to LLMs. Empirical results show that using selected triplets for fine-tuning leads to superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。