arXiv:2510.16769cs.AIcs.CL2025-10ACL

用视觉语言模型提升大规模图结构理解能力

See or Say Graphs: Agent-Driven Scalable Graph Structure Understanding with Vision-Language Models

  • 构建分层图检索框架,压缩冗余信息并保留关键推理内容
  • 支持200倍于现有基准的大规模图,性能比顶尖方法高4.4倍
  • 智能调度文本与视觉模态,发挥各自在属性和拓扑分析中的优势

视觉语言模型(VLMs)在图结构理解方面展现出潜力,但受限于输入令牌数量,面临可扩展性瓶颈,且缺乏有效机制协调文本与视觉模态。为此,我们提出GraphVista统一框架,同时提升可扩展性与模态协同能力。为解决可扩展性问题,GraphVista将图信息分层组织为轻量级GraphRAG基底,仅检索任务相关的文本描述和高分辨率视觉子图,压缩冗余上下文而保留关键推理元素。为增强模态协同,引入规划代理,将任务分解并路由至最合适的模态:使用文本模态直接访问显式图属性,利用视觉模态基于显式拓扑进行局部结构推理。大量实验表明,GraphVista可处理最大达现有基准200倍规模的图,持续优于现有纯文本、纯视觉及融合方法,在性能上最高实现4.4倍于当前最优基线的提升,充分挖掘双模态互补优势。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown promise in graph structure understanding, but remain limited by input-token constraints, facing scalability bottlenecks and lacking effective mechanisms to coordinate textual and visual modalities. To address these challenges, we propose GraphVista, a unified framework that enhances both scalability and modality coordination in graph structure understanding. For scalability, GraphVista organizes graph information hierarchically into a lightweight GraphRAG base, which retrieves only task-relevant textual descriptions and high-resolution visual subgraphs, compressing redundant context while preserving key reasoning elements. For modality coordination, GraphVista introduces a planning agent that decomposes and routes tasks to the most suitable modality-using the text modality for direct access to explicit graph properties and the visual modality for local graph structure reasoning grounded in explicit topology. Extensive experiments demonstrate that GraphVista scales to large graphs, up to 200$\times$ larger than those used in existing benchmarks, and consistently outperforms existing textual, visual, and fusion-based methods, achieving up to 4.4$\times$ quality improvement over the state-of-the-art baselines by fully exploiting the complementary strengths of both modalities.

图神经网络多模态可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。