arXiv:2607.06295cs.CV2026-07

研究图像图结构对分类性能的影响,发现结构设计显著决定模型效果。

Visual graphs for image classification: does the structure affect performance?

论文配图:Visual graphs for image classification: does the structure affect performance?
图 1 · 摘自论文原文
  • 固定三层GCN架构,系统比较多种图构建方法。
  • 不同图结构导致分类准确率差异明显,最高差达8.5%。
  • 适合关注视觉结构建模与图神经网络应用的研究者。

深度学习模型在各类视觉任务中表现出色,但难以充分编码图像中的内在视觉结构,常忽略空间、拓扑和语义信息。图神经网络为此提供了良好框架,但其在视觉任务中的有效应用仍不充分,且多从有限视角出发。本文通过固定三层数的GCN架构,系统比较当前主流的图构建技术。实证研究表明,网络结构显著影响性能,且图构建前的计算阶段也受结构本身强烈影响。该工作为视觉图神经网络的前期处理提供了重要方法论参考。

原文摘要 · Abstract (English)

Deep learning models have emerged in machine learning and related fields, demonstrating astonishing performance in various visual tasks. Despite their great success, however, these models are unable to fully encode intrinsic visual structures, and often ignore the spatial, topological, and semantic information contained within an image. Graph neural networks offer a good framework to face this aspect, but their effective use for visual tasks has been only partly explored and mainly starting from a limited perspective. This work aims to address this gap by conducting a systematic comparison of current graph construction techniques within the context of a fixed three-layer GCN architecture. Through an empirical study, it demonstrates in particular how the network structure affects performance and provides an important methodological contribution regarding the computational stages preceding graph utilization, which will be strongly influenced by the structure itself.

图神经网络图像分类结构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。