用细粒度多模态标记构建可跨知识图谱迁移的推理模型
Every Little Helps: Building Knowledge Graph Foundation Model with Fine-grained Transferable Multi-modal Tokens
- 将结构、视觉、文本信息转为专用标记,实现多模态融合
- 在17个知识图谱上均超越基线,对未见图谱表现优异
- 适合需要跨域推理的多模态知识图谱研究者
多模态知识图谱推理(MMKGR)旨在利用图结构和多模态实体内容预测缺失链接。现有方法多针对特定数据集,泛化能力弱。近期知识图谱基础模型(KGFMs)虽提升跨图谱迁移能力,但主要依赖结构模式,忽视丰富多模态信号。本文提出基于标记的础模型(TOFU),实现对不同多模态知识图谱的强泛化。TOFU将结构、视觉与文本信息离散化为模态特异标记,并采用分层融合架构与消息混合机制,处理这些标记以获取可迁移特征。在17个归纳性、直推式及完全归纳性多模态知识图谱上的实验表明,TOFU持续优于强基线模型,在未见知识图谱上表现突出。
原文摘要 · Abstract (English)
Multi-modal knowledge graph reasoning (MMKGR) aims to predict the missing links by exploiting both graph structure information and multi-modal entity contents. Most existing works are designed for a transductive setting, which learns dataset-specific embeddings and struggles to generalize to new KGs. Recent knowledge graph foundation models (KGFMs) improve cross-KG transfer, but they mainly exploit structural patterns and ignore rich multi-modal signals. We address these gaps by proposing a token-based foundation model (TOFU) for MMKGR, which exhibits strong generalization across different MMKGs. TOFU discretizes structural, visual, and textual information into modality-specific tokens. TOFU then employs a hierarchical fusion architecture with mixture-of-message mechanisms, aiming to process these tokens and obtain transferable features for MMKGR. Experimental results on 17 transductive, inductive, and fully-inductive MMKGs show that TOFU consistently outperforms strong KGFM and MMKGR baselines, delivering strong performance on unseen MMKGs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。