arXiv:2503.06974cs.CV2025-03AAAI被引 15

提出不对称视觉语义嵌入框架,提升图文匹配效率与精度

Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment

  • 根据文本输入动态选取图像不同区域特征,实现精准匹配
  • 通过径向偏置采样获取多视角图像特征,增强视觉信息表达
  • 基于语义单元匹配机制,显著提升跨模态相似度计算效率

学习视觉语义相似性是弥合图像与文本鸿沟的关键挑战。然而,视觉与语言数据存在固有差异,如信息密度不一——图像可能包含多个视角的文本信息,导致跨模态相似性计算难以准确高效。本文提出一种新型框架:不对称视觉语义嵌入(AVSE),可根据不同文本输入动态选择图像各区域特征以进行相似性计算。为捕捉图像中多视角信息,设计了径向偏置采样模块,用于采样图像块并提取多视角特征。此外,AVSE引入新模块,高效计算异构图像与文本嵌入间的视觉语义相似性。核心在于假设嵌入中存在基础语义单元,称为“元语义嵌入”,将其分割为同维元语义嵌入,并通过最优匹配方式计算两模态间的相似性。在大规模MS-COCO和Flickr30K数据集上广泛评估,结果表明本模型优于近期先进方法。

原文摘要 · Abstract (English)

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modalities accurately and efficiently. In this paper, we propose a novel framework called Asymmetric Visual Semantic Embedding (AVSE) to dynamically select features from various regions of images tailored to different textual inputs for similarity calculation. To capture information from different views in the image, we design a radial bias sampling module to sample image patches and obtain image features from various views, Furthermore, AVSE introduces a novel module for efficient computation of visual semantic similarity between asymmetric image and text embeddings. Central to this module is the presumption of foundational semantic units within the embeddings, denoted as ``meta-semantic embeddings." It segments all embeddings into meta-semantic embeddings with the same dimension and calculates visual semantic similarity by finding the optimal match of meta-semantic embeddings of two modalities. Our proposed AVSE model is extensively evaluated on the large-scale MS-COCO and Flickr30K datasets, demonstrating its superiority over recent state-of-the-art methods.

图文对齐视觉语义嵌入学习多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。