arXiv:2501.19069cs.CVcs.AI2025-01

用图脉冲网络提升视觉语言对齐,捕捉物体间复杂关系

Improving vision-language alignment with graph spiking hybrid Networks

  • 结合全景分割与图脉冲网络,生成细粒度语义特征
  • 在多个视觉语言任务上显著优于基线模型
  • 适合关注跨模态对齐与高效表征的科研人员

为弥合视觉与语言间的语义鸿沟,需设计有效的对齐策略,涵盖语义多样性、视觉信息的抽象表示及模型泛化能力。现有方法多依赖检测器框或规则划分的图像块来表征视觉语义,虽有进展,但仍难以充分捕捉不同物体间的细微上下文关系。本文提出一种综合视觉语义表示模块,依赖全景分割生成连贯的细粒度语义特征。同时,提出新型图脉冲混合网络(GSHN),融合脉冲神经网络(SNNs)与图注意力网络(GATs)的优势,编码视觉语义信息。该模型不仅能表示实例的离散与连续潜在变量,还能有效捕捉局部与全局上下文特征,显著增强语义表征的丰富性与多样性。利用SNN固有的时空特性,采用对比学习(CL)构建正负样本对,提升嵌入表示的相似性,降低计算开销并丰富有意义的视觉表征。设计创新的预训练方法——脉冲文本学习(STL),以文本特征增强离散语义编码能力。实验表明,所提GSHN在多个视觉语言下游任务中表现优异。

原文摘要 · Abstract (English)

To bridge the semantic gap between vision and language (VL), it is necessary to develop a good alignment strategy, which includes handling semantic diversity, abstract representation of visual information, and generalization ability of models. Recent works use detector-based bounding boxes or patches with regular partitions to represent visual semantics. While current paradigms have made strides, they are still insufficient for fully capturing the nuanced contextual relations among various objects. This paper proposes a comprehensive visual semantic representation module, necessitating the utilization of panoptic segmentation to generate coherent fine-grained semantic features. Furthermore, we propose a novel Graph Spiking Hybrid Network (GSHN) that integrates the complementary advantages of Spiking Neural Networks (SNNs) and Graph Attention Networks (GATs) to encode visual semantic information. Intriguingly, the model not only encodes the discrete and continuous latent variables of instances but also adeptly captures both local and global contextual features, thereby significantly enhancing the richness and diversity of semantic representations. Leveraging the spatiotemporal properties inherent in SNNs, we employ contrastive learning (CL) to enhance the similarity-based representation of embeddings. This strategy alleviates the computational overhead of the model and enriches meaningful visual representations by constructing positive and negative sample pairs. We design an innovative pre-training method, Spiked Text Learning (STL), which uses text features to improve the encoding ability of discrete semantics. Experiments show that the proposed GSHN exhibits promising results on multiple VL downstream tasks.

视觉语言对齐图神经网络脉冲神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。