arXiv:2504.08710cs.CV2025-04CVPR被引 21

用超图结构提升ViT对图像语义关系的建模能力。

Hypergraph Vision Transformers: Images are More than Nodes, More than Edges

  • 引入分层二部超图,无需聚类即可动态构建关系网络。
  • 在图像分类与检索任务中表现优异,计算效率更高。
  • 适合需要高效建模复杂语义关系的视觉任务研究者。

视觉变压器(ViTs)在计算机视觉中展现出良好可扩展性,但难以兼顾适应性、计算效率和高阶关系建模。视觉图神经网络(ViGs)虽提供替代方案,却受限于边生成中聚类算法的计算瓶颈。为此,我们提出超图视觉变压器(HgVT),将分层二部超图结构融入视觉变压器框架,以捕捉高阶语义关系并保持计算效率。HgVT采用群体与多样性正则化实现无聚类的动态超图构建,并通过专家边池化增强语义提取,支持基于图的图像检索。实验表明,HgVT在图像分类与检索任务上均取得优异性能,是语义驱动视觉任务的高效框架。

原文摘要 · Abstract (English)

Recent advancements in computer vision have highlighted the scalability of Vision Transformers (ViTs) across various tasks, yet challenges remain in balancing adaptability, computational efficiency, and the ability to model higher-order relationships. Vision Graph Neural Networks (ViGs) offer an alternative by leveraging graph-based methodologies but are hindered by the computational bottlenecks of clustering algorithms used for edge generation. To address these issues, we propose the Hypergraph Vision Transformer (HgVT), which incorporates a hierarchical bipartite hypergraph structure into the vision transformer framework to capture higher-order semantic relationships while maintaining computational efficiency. HgVT leverages population and diversity regularization for dynamic hypergraph construction without clustering, and expert edge pooling to enhance semantic extraction and facilitate graph-based image retrieval. Empirical results demonstrate that HgVT achieves strong performance on image classification and retrieval, positioning it as an efficient framework for semantic-based vision tasks.

视觉变压器超图图像检索语义建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。