arXiv:2501.18119cs.CLcs.AI2025-01ACL被引 21

用离散代码让大模型无缝理解知识图谱结构

Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models

  • 自监督学习将图谱结构与语义压缩为离散编码
  • 每实体仅需16个令牌,性能超越传统提示方法
  • 适合想融合知识图谱的大模型应用者

由于知识图谱(KG)结构与自然语言之间存在天然鸿沟,如何有效整合KG的全局结构信息与大语言模型(LLMs)成为关键问题。为此,我们提出一种两阶段框架,通过学习并应用每个实体的量化编码,实现知识图谱与大模型的无缝融合。首先,提出自监督量化表示(SSQR)方法,将KG的结构与语义知识压缩为离散代码(即标记),使其格式与自然语言句式对齐;进一步设计基于这些编码的图谱指令跟随数据,直接输入大模型,实现无缝集成。实验表明,SSQR优于现有无监督量化方法,生成更具区分度的编码;微调后的LLaMA2和LLaMA3.1在知识图谱链接预测和三元组分类任务中表现更优,且每实体仅需16个令牌,远少于传统提示方法中的数千个。

原文摘要 · Abstract (English)

Due to the presence of the natural gap between Knowledge Graph (KG) structures and the natural language, the effective integration of holistic structural information of KGs with Large Language Models (LLMs) has emerged as a significant question. To this end, we propose a two-stage framework to learn and apply quantized codes for each entity, aiming for the seamless integration of KGs with LLMs. Firstly, a self-supervised quantized representation (SSQR) method is proposed to compress both KG structural and semantic knowledge into discrete codes (\ie, tokens) that align the format of language sentences. We further design KG instruction-following data by viewing these learned codes as features to directly input to LLMs, thereby achieving seamless integration. The experiment results demonstrate that SSQR outperforms existing unsupervised quantized methods, producing more distinguishable codes. Further, the fine-tuned LLaMA2 and LLaMA3.1 also have superior performance on KG link prediction and triple classification tasks, utilizing only 16 tokens per entity instead of thousands in conventional prompting methods.

知识图谱大模型融合量化表示自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。