arXiv:2505.07895cs.LGcs.AI2025-05IJCAI被引 1

通过模态间相互影响提升异构网络节点分类效果

Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks

  • 引入嵌套模态注意力机制,动态融合多模态信息
  • 在异构图上实现跨模态一致性传播,提升表示质量
  • 适合处理带网络结构的多模态数据,尤其缺模态场景

如今,众多在线平台可被建模为多模态异构网络(MMHNs),如豆瓣电影网络和亚马逊产品评论网络。准确对这些网络中的节点进行分类对于分析对应实体至关重要,这需要在节点上进行有效的表示学习。然而,现有方法通常采用早期融合策略,可能丢失各模态的独特特征;或采用晚期融合方式,忽略基于GNN的信息传播中跨模态的引导作用。本文提出一种新型模型——异构图神经网络带模态间注意力(HGNN-IMA),通过在信息传播过程中捕捉多模态间的相互影响来学习节点表示,框架基于异构图变压器。具体而言,将嵌套的模态间注意力机制融入节点间注意力,实现自适应多模态融合,并考虑模态对齐,以促进所有模态下相似性一致的节点间传播。此外,引入注意力损失以缓解缺失模态的影响。大量实验验证了该模型在节点分类任务上的优越性,为处理多模态数据(尤其是带有网络结构的数据)提供了新视角。

原文摘要 · Abstract (English)

Nowadays, numerous online platforms can be described as multi-modal heterogeneous networks (MMHNs), such as Douban's movie networks and Amazon's product review networks. Accurately categorizing nodes within these networks is crucial for analyzing the corresponding entities, which requires effective representation learning on nodes. However, existing multi-modal fusion methods often adopt either early fusion strategies which may lose the unique characteristics of individual modalities, or late fusion approaches overlooking the cross-modal guidance in GNN-based information propagation. In this paper, we propose a novel model for node classification in MMHNs, named Heterogeneous Graph Neural Network with Inter-Modal Attention (HGNN-IMA). It learns node representations by capturing the mutual influence of multiple modalities during the information propagation process, within the framework of heterogeneous graph transformer. Specifically, a nested inter-modal attention mechanism is integrated into the inter-node attention to achieve adaptive multi-modal fusion, and modality alignment is also taken into account to encourage the propagation among nodes with consistent similarities across all modalities. Moreover, an attention loss is augmented to mitigate the impact of missing modalities. Extensive experiments validate the superiority of the model in the node classification task, providing an innovative view to handle multi-modal data, especially when accompanied with network structures.

图神经网络多模态学习节点分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。