arXiv:2601.21453cs.AI2026-01被引 2

用克利福德代数实现多模态图的精准对齐与自适应融合

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

  • 基于克利福德代数构建多模态几何流形,实现高阶图传播与模态交互
  • 提出自适应全息聚合模块,利用几何等级特性提升融合效果
  • 在9个数据集上超越当前最优模型,适用于多模态图学习任务

近年来,多模态领域的快速发展推动了图机器学习向以数据为中心的范式转变,从文本属性图扩展到多模态属性图。这一进展显著增强了数据表征能力,拓展了下游任务范围,如模态导向任务,提升了图学习的实际应用价值。然而现有神经范式存在两大局限:(1) 模态对齐中忽略图上下文,多数方法采用拓扑约束或模态特定算子作为分词器,导致模态间交互受限;(2) 模态融合缺乏自适应性,多数方法仅适用于双模态场景,未能充分利用带有拓扑先验的对齐特征,造成泛化能力差、性能下降。为此,我们提出LION(Clifford Neural Paradigm),基于克利福德代数与解耦的图神经网络范式(即传播-聚合),实现多模态属性图中的对齐-融合流程。具体而言,首先构建基于克利福德代数的模态感知几何流形,通过几何诱导的高阶图传播实现模态交互,促进模态对齐;随后,基于对齐特征的几何等级特性,提出自适应全息聚合模块,将几何等级的能量与尺度结合可学习参数,优化模态融合。在9个数据集上的大量实验表明,LION在3类图任务和3类模态下游任务中显著优于当前最优基线。

原文摘要 · Abstract (English)

Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data representation and expands the scope of graph downstream tasks, such as modality-oriented tasks, thereby improving the practical utility of graph ML. Despite its promise, limitations exist in the current neural paradigms: (1) Neglect Context in Modality Alignment: Most existing methods adopt topology-constrained or modality-specific operators as tokenizers. These aligners inevitably neglect graph context and inhibit modality interaction, resulting in suboptimal alignment. (2) Lack of Adaptation in Modality Fusion: Most existing methods are simple adaptations for 2-modality graphs and fail to adequately exploit aligned tokens equipped with topology priors during fusion, leading to poor generalizability and performance degradation. To address the above issues, we propose LION (c\underline{LI}ff\underline{O}rd \underline{N}eural paradigm) based on the Clifford algebra and decoupled graph neural paradigm (i.e., propagation-then-aggregation) to implement alignment-then-fusion in multimodal-attributed graphs. Specifically, we first construct a modality-aware geometric manifold grounded in Clifford algebra. This geometric-induced high-order graph propagation efficiently achieves modality interaction, facilitating modality alignment. Then, based on the geometric grade properties of aligned tokens, we propose adaptive holographic aggregation. This module integrates the energy and scale of geometric grades with learnable parameters to improve modality fusion. Extensive experiments on 9 datasets demonstrate that LION significantly outperforms SOTA baselines across 3 graph and 3 modality downstream tasks.

多模态图克利福德代数图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。