arXiv:2510.24788cs.CVcs.AI2025-10NeurIPS被引 7

视觉模型能像人一样看懂图的全局结构,表现优于传统图神经网络。

The Underappreciated Power of Vision Models for Graph Structural Understanding

  • 用视觉模型捕捉图的全局结构,不依赖逐层消息传递。
  • 在新基准上,视觉模型在识别对称性和关键节点上远超图神经网络。
  • 适合需要整体感知和跨尺度推理的任务,如复杂网络分析。

图神经网络通过自下而上的消息传递运行,与人类直观先感知整体结构的方式截然不同。我们研究了视觉模型在图理解中的潜在能力,发现其在现有基准上性能可媲美图神经网络,但学习模式显著不同。由于现有基准混淆了领域特征与拓扑理解,我们提出GraphAbstract,用于评估模型像人类一样感知全局图属性的能力:识别组织原型、检测对称性、感知连通强度、定位关键元素。结果表明,视觉模型在需要整体结构理解的任务中显著优于图神经网络,并在不同图规模间保持泛化能力,而图神经网络在全局模式抽象上表现不佳,且随图变大性能下降。本工作证明,视觉模型在图结构理解方面具有未被重视的强大能力,尤其适用于需全局拓扑意识和尺度不变推理的问题。这些发现为构建更有效的图基础模型开辟了新路径。

原文摘要 · Abstract (English)

Graph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visual perception, which intuitively captures global structures first. We investigate the underappreciated potential of vision models for graph understanding, finding they achieve performance comparable to GNNs on established benchmarks while exhibiting distinctly different learning patterns. These divergent behaviors, combined with limitations of existing benchmarks that conflate domain features with topological understanding, motivate our introduction of GraphAbstract. This benchmark evaluates models' ability to perceive global graph properties as humans do: recognizing organizational archetypes, detecting symmetry, sensing connectivity strength, and identifying critical elements. Our results reveal that vision models significantly outperform GNNs on tasks requiring holistic structural understanding and maintain generalizability across varying graph scales, while GNNs struggle with global pattern abstraction and degrade with increasing graph size. This work demonstrates that vision models possess remarkable yet underutilized capabilities for graph structural understanding, particularly for problems requiring global topological awareness and scale-invariant reasoning. These findings open new avenues to leverage this underappreciated potential for developing more effective graph foundation models for tasks dominated by holistic pattern recognition.

视觉模型图理解结构感知新基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。