arXiv:2511.07710cs.CVcs.MM2025-11AAAI被引 1

通过细粒度建模与区域不确定性,提升图像与文本的精准对应能力。

Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling

  • 用模态特异性偏差识别关键视觉与文本特征,避免依赖脆弱的跨模态注意力。
  • 将区域特征建模为高斯混合分布,捕捉一对多、多对一的复杂对应关系。
  • 在Flickr30K和MS-COCO上表现领先,适合视觉问答与图像描述等任务。

细粒度图像-文本对齐是多模态学习中的核心挑战,支撑视觉问答、图像字幕生成和视觉语言导航等应用。与全局对齐不同,细粒度对齐需实现局部视觉区域与文本词元间的精确对应,常受噪声注意力机制和跨模态关系简化建模的阻碍。本文指出现有方法存在两大缺陷:缺乏鲁棒的模内机制评估视觉与文本词元的重要性,导致复杂场景下泛化能力差;且缺乏细粒度不确定性建模,无法捕捉区域-词的一对多与多对一关系。为此,我们提出统一方法,融合显著性感知与粒度感知建模,以及区域级不确定性建模。该方法利用模态特异性偏差识别显著特征,不依赖脆弱的跨模态注意力;并将区域特征表示为高斯混合分布,以捕捉细粒度不确定性。在Flickr30K和MS-COCO上的大量实验表明,该方法在多种主干网络上均达当前最优性能,显著提升了细粒度图像-文本对齐的鲁棒性与可解释性。

原文摘要 · Abstract (English)

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual regions and textual tokens, often hindered by noisy attention mechanisms and oversimplified modeling of cross-modal relationships. In this work, we identify two fundamental limitations of existing approaches: the lack of robust intra-modal mechanisms to assess the significance of visual and textual tokens, leading to poor generalization in complex scenes; and the absence of fine-grained uncertainty modeling, which fails to capture the one-to-many and many-to-one nature of region-word correspondences. To address these issues, we propose a unified approach that incorporates significance-aware and granularity-aware modeling and region-level uncertainty modeling. Our method leverages modality-specific biases to identify salient features without relying on brittle cross-modal attention, and represents region features as a mixture of Gaussian distributions to capture fine-grained uncertainty. Extensive experiments on Flickr30K and MS-COCO demonstrate that our approach achieves state-of-the-art performance across various backbone architectures, significantly enhancing the robustness and interpretability of fine-grained image-text alignment.

多模态对齐细粒度匹配不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。