arXiv:2512.17178cs.CVcs.IR2025-12

无需训练,提升图像文本匹配中属性与物体的精准绑定。

ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching

  • 通过语义精炼机制优化文本中物体和属性的词元表示
  • 利用局部图文对齐计算相似度,显著提升绑定准确率
  • 完全免训练,适用于新组合概念,适合快速部署

对比语言-图像预训练(CLIP)在多种多模态任务中表现优异,但在组合式图像-文本匹配上仍面临挑战,尤其难以准确关联物体与其属性,因其全局表征常忽略细粒度语义。现有方法通常需要额外训练或大量难负样本,但泛化能力有限,且未根本解决全局表征的缺陷。本文提出ABE-CLIP,一种无需训练的属性绑定增强方法,旨在强化类CLIP模型中的属性-物体绑定。我们设计语义精炼机制,对文本中物体和属性短语的词元嵌入进行优化,减轻属性混淆,提升语义精度;进一步引入局部词元-图像块对齐策略,计算经精炼的文本词元与最相关图像块间的相似度,并通过聚合局部相似度得到最终图像-文本相似度。在多个数据集上的实验表明,ABE-CLIP显著提升了属性-物体绑定性能,甚至超越需大量训练的方法。

原文摘要 · Abstract (English)

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable performance in various multimodal tasks. However, it still struggles with compositional image-text matching, particularly in accurately associating objects with their corresponding attributes, because its inherent global representation often overlooks fine-grained semantics for attribute binding. Existing methods often require additional training or extensive hard negative sampling, yet they frequently show limited generalization to novel compositional concepts and fail to fundamentally address the drawbacks of global representations. In this paper, we propose ABE-CLIP, a novel training-free Attribute Binding Enhancement method designed to strengthen attribute-object binding in CLIP-like models. Specifically, we employ a Semantic Refinement Mechanism to refine token embeddings for both object and attribute phrases in the text, thereby mitigating attribute confusion and improving semantic precision. We further introduce a Local Token-Patch Alignment strategy that computes similarity scores between refined textual tokens and their most relevant image patches. By aggregating localized similarity scores, ABE-CLIP computes the final image-text similarity. Experiments on multiple datasets demonstrate that ABE-CLIP significantly improves attribute-object binding performance, even surpassing methods that require extensive training.

图像文本匹配属性绑定免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。