arXiv:2608.24795cs.LG2026-08被引 2

提出LION框架,用克利福德代数实现多模态图的精准对齐与自适应融合。

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

论文配图:LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning
图 1 · 摘自论文原文
  • 基于克利福德代数构建模态感知几何流形,实现高阶图传播与模态交互。
  • 在9个图文多模态图数据集上,3项图任务和3项模态任务均超越现有最优方法。
  • 适合需要多模态图表示学习与跨模态融合的科研与工业场景。

近年来,多模态领域的快速发展推动了图机器学习向以数据为中心的范式转变,从文本属性图扩展到多模态属性图。这一进展显著增强了数据表征能力,并拓展了图下游任务(如模态导向任务)的应用范围,提升了图学习的实际价值。然而,现有神经范式仍存在两大局限:(1) 模态对齐中忽略图上下文:多数方法采用拓扑约束或模态专用算子作为分词器,导致对齐过程忽略图结构信息,抑制模态间交互,造成对齐效果不佳;(2) 模态融合缺乏适应性:现有方法多为双模态图的简单适配,未能有效利用具有拓扑先验的对齐特征进行融合,导致泛化能力差、性能下降。为此,我们提出LION(Clifford Neural Paradigm),基于克利福德代数与解耦图神经网络范式(即传播-聚合分离),实现多模态属性图中的对齐-融合流程。具体而言,首先构建基于克利福德代数的模态感知几何流形,通过该几何诱导的高阶图传播实现模态间高效交互,促进模态对齐。随后,基于对齐后令牌的拓扑感知克利福德分量,提出自适应全息聚合模块,该模块结合分量能量与传播尺度信息,引入可学习参数以增强模态融合。在9个文本-图像多模态图(MAG)数据集上的大量实验表明,LION在3类图任务和3类模态下游任务中均显著优于现有最先进基线方法。

原文摘要 · Abstract (English)

Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data representation and expands the scope of graph downstream tasks, such as modality-oriented tasks, thereby improving the practical utility of graph ML. Despite its promise, limitations exist in the current neural paradigms:(1) Neglect Context in Modality Alignment: Most existing methods adopt topology-constrained or modality-specific operators as tokenizers.These aligners inevitably neglect graph context and inhibit modality interaction, resulting in suboptimal alignment.(2) Lack of Adaptation in Modality Fusion: Most existing methods are simple adaptations for 2-modality graphs and fail to adequately exploit aligned tokens equipped with topology priors during fusion, leading to poor generalizability and performance degradation.To address the above issues, we propose LION (c\underline{LI}ff\underline{O}rd \underline{N}eural paradigm) based on the Clifford algebra and decoupled graph neural paradigm (i.e., propagation-then-aggregation) to implement alignment-then-fusion in multimodal-attributed graphs. Specifically, we first construct a modality-aware geometric manifold grounded in Clifford algebra.This geometric-induced high-order graph propagation efficiently achieves modality interaction, facilitating modality alignment.Then, based on the topology-aware Clifford components of aligned tokens, we propose adaptive holographic aggregation. This module integrates component-wise energy and propagation-scale information with learnable parameters to improve modality fusion. Extensive experiments on 9 text-image MAG datasets demonstrate that LION significantly outperforms SOTA baselines across 3 graph and 3 modality downstream tasks.

多模态图克利福德代数模态对齐图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。