用超复数统一融合与独立表示,提升多模态知识图谱补全性能。
Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion
- 引入双四元数框架,将四种模态表示映射到正交基上
- 在三个基准数据集上达到最优效果,比现有方法提升2.3%~4.7%
- 适合需要兼顾模态独立性与跨模态交互的场景
多模态知识图谱补全(MMKGC)旨在利用实体的结构关系和多源模态信息,发现多模态知识图谱中的缺失事实。现有方法分为融合式与集成式:前者采用固定融合策略,导致模态特异性信息丢失;后者虽保留模态独立性,却难以捕捉模态间的上下文依赖语义交互。为此,本文提出M-Hyper方法,实现融合与独立模态表示的协同共存。受四元数代数启发,利用其四个正交基表示多个独立模态,并通过哈密顿积高效建模模态间两两交互。设计细粒度实体表示分解(FERF)模块与鲁棒关系感知模态融合(R2MF)模块,分别获取三种独立模态和一种融合模态的鲁棒表示。最终,四种模态表示被映射至双四元数的四个正交基,实现全面的模态交互。大量实验表明,该方法在三个基准数据集上均取得最先进性能,兼具鲁棒性与计算效率。
原文摘要 · Abstract (English)
Multi-modal knowledge graph completion (MMKGC) aims to discover missing facts in multi-modal knowledge graphs (MMKGs) by leveraging both structural relationships and diverse modality information of entities. Existing MMKGC methods follow two multi-modal paradigms: fusion-based and ensemble-based. Fusion-based methods employ fixed fusion strategies, which inevitably leads to the loss of modality-specific information and a lack of flexibility to adapt to varying modality relevance across contexts. In contrast, ensemble-based methods retain modality independence through dedicated sub-models but struggle to capture the nuanced, context-dependent semantic interplay between modalities. To overcome these dual limitations, we propose a novel MMKGC method M-Hyper, which achieves the coexistence and collaboration of fused and independent modality representations. Our method integrates the strengths of both paradigms, enabling effective cross-modal interactions while maintaining modality-specific information. Inspired by ``quaternion'' algebra, we utilize its four orthogonal bases to represent multiple independent modalities and employ the Hamilton product to efficiently model pair-wise interactions among them. Specifically, we introduce a Fine-grained Entity Representation Factorization (FERF) module and a Robust Relation-aware Modality Fusion (R2MF) module to obtain robust representations for three independent modalities and one fused modality. The resulting four modality representations are then mapped to the four orthogonal bases of a biquaternion (a hypercomplex extension of quaternion) for comprehensive modality interaction. Extensive experiments indicate its state-of-the-art performance, robustness, and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。