arXiv:2603.11520cs.CVcs.AI2026-03

解决图像检索中视觉与文本注意力失衡问题,提升复杂场景下的检索准确率。

FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval

  • 提出FBCIR方法,识别影响检索决策的关键视觉与文本成分。
  • 在硬负样本场景下,现有模型准确率显著下降,存在明显模态注意力失衡。
  • 设计数据增强流程,引入精心构造的难例负样本,促进跨模态平衡推理。

组合图像检索(CIR)要求多模态模型对文本-图像输入对中的视觉内容和语义修改进行联合推理。尽管当前CIR模型在常见基准测试中表现优异,但在更具挑战性的场景下性能常显著下降,尤其是当负样本在语义上与查询图像或文本高度相关时。本文将此现象归因于注意力失衡,即模型过度关注某一模态而忽略另一模态。为此,我们提出FBCIR——一种多模态焦点解释方法,用于识别决定模型检索决策的关键视觉与文本组件。通过FBCIR分析发现,现有CIR模型在困难负样本设置下普遍存在注意力失衡。基于此分析,我们进一步设计了一种数据增强工作流,为现有CIR数据集添加经筛选的难例负样本,以促进更均衡的跨模态推理。在多个CIR模型上的大量实验表明,该增强方法在复杂场景下持续提升性能,同时保持在标准基准上的能力。FBCIR解释方法与数据增强工作流共同为CIR模型诊断与鲁棒性提升提供了新视角。

原文摘要 · Abstract (English)

Composed image retrieval (CIR) requires multi-modal models to jointly reason over visual content and semantic modifications presented in text-image input pairs. While current CIR models achieve strong performance on common benchmark cases, their accuracies often degrades in more challenging scenarios where negative candidates are semantically aligned with the query image or text. In this paper, we attribute this degradation to focus imbalances, where models disproportionately attend to one modality while neglecting the other. To validate this claim, we propose FBCIR, a multi-modal focus interpretation method that identifies the most crucial visual and textual input components to a model's retrieval decisions. Using FBCIR, we report that focus imbalances are prevalent in existing CIR models, especially under hard negative settings. Building on the analyses, we further propose a CIR data augmentation workflow that facilitates existing CIR datasets with curated hard negatives designed to encourage balanced cross-modal reasoning. Extensive experiments across multiple CIR models demonstrate that the proposed augmentation consistently improves performance in challenging cases, while maintaining their capabilities on standard benchmarks. Together, our interpretation method and data augmentation workflow provide a new perspective on CIR model diagnosis and robustness improvements.

图像检索多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。