arXiv:2503.21435cs.AI2025-03被引 6

用视觉语言模型解决多图联合理解与推理,填补了该领域空白。

Graph-to-Vision: Multi-graph Understanding and Reasoning using Vision-Language Models

  • 构建首个多图联合推理评测基准,覆盖四类常见图结构。
  • 在多维度评估中验证主流视觉语言模型的推理能力并提升效果。
  • 适合对跨模态图智能、多图推理感兴趣的学者和开发者。

视觉语言模型(VLMs)在解析可视化图数据方面展现出潜力,为超越传统图神经网络(GNNs)的图结构推理提供了新视角。然而,现有研究主要集中于单图推理,多图联合推理仍缺乏探索。本文提出首个全面的多图推理评测基准,涵盖知识图谱、流程图、思维导图和路线图四类常见图类型,支持同质与异质图组,并包含复杂度递增的任务。我们采用多维评分框架评估多个前沿VLMs在图解析、推理一致性和指令遵循准确性上的表现。同时,微调多个开源模型后观察到性能持续提升,验证了数据集的有效性。本工作为推进多图理解提供了系统性路径,揭示了跨模态图智能的新机遇。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have shown promising capabilities in interpreting visualized graph data, offering a new perspective for graph-structured reasoning beyond traditional Graph Neural Networks (GNNs). However, existing studies focus primarily on single-graph reasoning, leaving the critical challenge of multi-graph joint reasoning underexplored. In this work, we introduce the first comprehensive benchmark designed to evaluate and enhance the multi-graph reasoning abilities of VLMs. Our benchmark covers four common graph types-knowledge graphs, flowcharts, mind maps, and route maps-and supports both homogeneous and heterogeneous graph groupings with tasks of increasing complexity. We evaluate several state-of-the-art VLMs under a multi-dimensional scoring framework that assesses graph parsing, reasoning consistency, and instruction-following accuracy. Additionally, we fine-tune multiple open-source models and observe consistent improvements, confirming the effectiveness of our dataset. This work provides a principled step toward advancing multi-graph understanding and reveals new opportunities for cross-modal graph intelligence.

多图推理视觉语言模型跨模态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。