arXiv:2506.02568cs.AI2025-06KDD被引 11

让大模型同时理解图文混合的复杂图结构,提升多模态图推理能力。

MLaGA: Multimodal Large Language and Graph Assistant

  • 设计统一编码器,对齐图文属性到同一空间。
  • 通过轻量投影器融合多模态特征与图结构,提升推理效果。
  • 适用于真实场景中的图文混合图分析,适合做多模态图学习的研究者。

大型语言模型(LLMs)在推进图结构数据的分析方面表现出显著成效。现有基于LLM的图方法主要针对文本丰富的图,其中节点属性为文本描述。然而,对于包含多种属性类型(如文本和图像)的多模态图,其应用仍处于探索阶段,尽管这类图在现实场景中极为普遍。为弥合这一差距,我们提出多模态大语言与图助手(MLaGA),一种创新模型,可有效拓展LLM能力,以实现对复杂图结构和多模态属性的推理。首先,我们设计了一个结构感知的多模态编码器,通过联合图预训练目标,将文本与视觉属性对齐至统一空间。随后,我们采用多模态指令微调方法,通过轻量级投影器,无缝将多模态特征与图结构融入到LLM中。在多个数据集上的大量实验表明,相比领先基线方法,MLaGA在监督和迁移学习场景下的各类图学习任务中均表现出更优性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated substantial efficacy in advancing graph-structured data analysis. Prevailing LLM-based graph methods excel in adapting LLMs to text-rich graphs, wherein node attributes are text descriptions. However, their applications to multimodal graphs--where nodes are associated with diverse attribute types, such as texts and images--remain underexplored, despite their ubiquity in real-world scenarios. To bridge the gap, we introduce the Multimodal Large Language and Graph Assistant (MLaGA), an innovative model that adeptly extends LLM capabilities to facilitate reasoning over complex graph structures and multimodal attributes. We first design a structure-aware multimodal encoder to align textual and visual attributes within a unified space through a joint graph pre-training objective. Subsequently, we implement a multimodal instruction-tuning approach to seamlessly integrate multimodal features and graph structures into the LLM through lightweight projectors. Extensive experiments across multiple datasets demonstrate the effectiveness of MLaGA compared to leading baseline methods, achieving superior performance in diverse graph learning tasks under both supervised and transfer learning scenarios.

多模态图神经网络大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。