在视觉语言联合空间中生成编辑,实现零样本图像组合检索
Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
- 在联合视觉语言空间中融合多模态特征进行编辑
- 仅用20万合成数据微调即达领先性能
- 适合需要高效零样本检索的研究者
组合图像检索(CIR)通过结合参考图像与文本修改实现细粒度视觉搜索。尽管监督方法准确率高,但依赖昂贵的三元组标注,推动了零样本解决方案的发展。零样本CIR的核心挑战在于:现有以文本为中心或基于扩散的方法难以有效弥合视觉-语言模态差距。为此,我们提出Fusion-Diff,一种高效且数据高效的生成编辑框架,用于多模态对齐。首先,在联合视觉语言(VL)空间中引入多模态融合特征编辑策略,显著缩小模态差距;其次,为最大化数据效率,框架集成轻量级Control-Adapter,仅在20万样本的合成数据集上微调即可达到当前最优性能。在标准CIR基准(CIRR、FashionIQ和CIRCO)上的大量实验表明,Fusion-Diff显著优于以往零样本方法。我们还通过可视化融合的多模态表示增强模型可解释性。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) enables fine-grained visual search by combining a reference image with a textual modification. While supervised CIR methods achieve high accuracy, their reliance on costly triplet annotations motivates zero-shot solutions. The core challenge in zero-shot CIR (ZS-CIR) stems from a fundamental dilemma: existing text-centric or diffusion-based approaches struggle to effectively bridge the vision-language modality gap. To address this, we propose Fusion-Diff, a novel generative editing framework with high effectiveness and data efficiency designed for multimodal alignment. First, it introduces a multimodal fusion feature editing strategy within a joint vision-language (VL) space, substantially narrowing the modality gap. Second, to maximize data efficiency, the framework incorporates a lightweight Control-Adapter, enabling state-of-the-art performance through fine-tuning on only a limited-scale synthetic dataset of 200K samples. Extensive experiments on standard CIR benchmarks (CIRR, FashionIQ, and CIRCO) demonstrate that Fusion-Diff significantly outperforms prior zero-shot approaches. We further enhance the interpretability of our model by visualizing the fused multimodal representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。