通过双分支结构增强细节,提升图像检索的精准度。
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
- 双分支设计分离全局与细节特征,动态融合
- 在CIRR和FashionIQ上达到最优性能,显著提升细粒度匹配能力
- 适合需要精确视觉修改理解的应用场景
组合图像检索(CIR)旨在根据参考图像和修改文本组成的联合查询,在图库中检索目标图像。现有方法侧重于跨模态全局信息的平衡编码,但因忽视细粒度细节,难以处理细微视觉变化或复杂文本指令。本文提出DetailFusion,一种新颖的双分支框架,有效协调全局与细节层次的信息,实现细节增强的CIR。方法利用来自图像编辑数据集的原子级细节变化先验,并结合面向细节的优化策略,构建细节感知推理分支;同时设计自适应特征复合器,根据每个多模态查询的细粒度特征动态融合全局与细节特征。大量实验与消融分析表明,该方法在CIRR和FashionIQ数据集上均达当前最优性能,验证了细节增强在CIR中的有效性及跨领域适应性。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) aims to retrieve target images from a gallery based on a reference image and modification text as a combined query. Recent approaches focus on balancing global information from two modalities and encode the query into a unified feature for retrieval. However, due to insufficient attention to fine-grained details, these coarse fusion methods often struggle with handling subtle visual alterations or intricate textual instructions. In this work, we propose DetailFusion, a novel dual-branch framework that effectively coordinates information across global and detailed granularities, thereby enabling detail-enhanced CIR. Our approach leverages atomic detail variation priors derived from an image editing dataset, supplemented by a detail-oriented optimization strategy to develop a Detail-oriented Inference Branch. Furthermore, we design an Adaptive Feature Compositor that dynamically fuses global and detailed features based on fine-grained information of each unique multimodal query. Extensive experiments and ablation analyses not only demonstrate that our method achieves state-of-the-art performance on both CIRR and FashionIQ datasets but also validate the effectiveness and cross-domain adaptability of detail enhancement for CIR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。