arXiv:2505.20422cs.CLcs.AI2025-05EMNLP被引 8

让知识图谱模型同时理解结构和文字,提升未知关系推理能力

SEMMA: A Semantic Aware Knowledge Graph Foundation Model

  • 用大模型增强关系语义,构建文本关系图与结构图融合
  • 在54个知识图谱上超越纯结构模型,未知关系预测效果翻倍
  • 适合需要跨领域泛化推理的研究者和应用开发

知识图谱基础模型(KGFMs)通过学习可迁移模式,在未见图上实现零样本推理方面展现出潜力。然而,现有多数KGFMs仅依赖图结构,忽视了文本属性中蕴含的丰富语义信息。我们提出SEMMA,一种双模块知识图谱基础模型,系统性地将可迁移的文本语义与图结构结合。SEMMA利用大语言模型(LLMs)丰富关系标识符,生成语义嵌入,并据此构建文本关系图,再与结构组件融合。在54个不同知识图谱上,SEMMA在全归纳链接预测任务中优于纯结构基线模型ULTRA。尤为重要的是,在更具挑战性的泛化场景下(测试时关系词汇完全未见过),结构方法性能崩溃,而SEMMA的有效性高出2倍。研究结果表明,当仅靠结构无法泛化时,文本语义至关重要,凸显了统一结构与语言信号的知识推理基础模型的必要性。

原文摘要 · Abstract (English)

Knowledge Graph Foundation Models (KGFMs) have shown promise in enabling zero-shot reasoning over unseen graphs by learning transferable patterns. However, most existing KGFMs rely solely on graph structure, overlooking the rich semantic signals encoded in textual attributes. We introduce SEMMA, a dual-module KGFM that systematically integrates transferable textual semantics alongside structure. SEMMA leverages Large Language Models (LLMs) to enrich relation identifiers, generating semantic embeddings that subsequently form a textual relation graph, which is fused with the structural component. Across 54 diverse KGs, SEMMA outperforms purely structural baselines like ULTRA in fully inductive link prediction. Crucially, we show that in more challenging generalization settings, where the test-time relation vocabulary is entirely unseen, structural methods collapse while SEMMA is 2x more effective. Our findings demonstrate that textual semantics are critical for generalization in settings where structure alone fails, highlighting the need for foundation models that unify structural and linguistic signals in knowledge reasoning.

知识图谱大模型语义融合零样本推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。