arXiv:2506.09738cs.LG2025-06被引 5

构建通用多模态图大模型,实现跨数据与任务的统一建模。

Towards Multimodal Graph Large Language Model

  • 提出多模态图统一框架,支持多粒度多尺度特征融合。
  • 定义五项核心能力,涵盖任务泛化与自然语言交互。
  • 适合图学习与多模态研究者,推动通用模型发展。

多模态图在现实应用中广泛存在,但现有方法通常针对特定图数据和任务从头训练,难以跨数据与任务泛化。为此,我们探索多模态图大语言模型(MG-LLM)的潜力,旨在统一并泛化于多样化的多模态图数据与任务。我们提出一个多模态图数据、任务与模型的统一框架,揭示其内在的多粒度、多尺度特性。具体地,提出五项关键能力:1)多模态结构与属性的统一表示空间;2)处理多样化多模态图任务的能力;3)多模态图上下文学习;4)与自然语言的交互能力;5)多模态图推理能力。随后分析关键技术挑战,综述相关工作,并指出未来研究方向。最后,总结了可用于模型训练的现有多模态图数据集。本文有助于推动MG-LLM在跨多模态图数据与任务中的通用性研究进展。

原文摘要 · Abstract (English)

Multi-modal graphs, which integrate diverse multi-modal features and relations, are ubiquitous in real-world applications. However, existing multi-modal graph learning methods are typically trained from scratch for specific graph data and tasks, failing to generalize across various multi-modal graph data and tasks. To bridge this gap, we explore the potential of Multi-modal Graph Large Language Models (MG-LLM) to unify and generalize across diverse multi-modal graph data and tasks. We propose a unified framework of multi-modal graph data, task, and model, discovering the inherent multi-granularity and multi-scale characteristics in multi-modal graphs. Specifically, we present five key desired characteristics for MG-LLM: 1) unified space for multi-modal structures and attributes, 2) capability of handling diverse multi-modal graph tasks, 3) multi-modal graph in-context learning, 4) multi-modal graph interaction with natural language, and 5) multi-modal graph reasoning. We then elaborate on the key challenges, review related works, and highlight promising future research directions towards realizing these ambitious characteristics. Finally, we summarize existing multi-modal graph datasets pertinent for model training. We believe this paper can contribute to the ongoing advancement of the research towards MG-LLM for generalization across multi-modal graph data and tasks.

多模态图大模型图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。