让图像检索支持多步文本修改,更贴近真实使用场景。
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval

- 提出Text-oriented Entity Mapping架构,精准对齐文本与图像实体
- 在四个数据集上验证,多修改场景下检索准确率显著提升
- 构建两个高质量多修改数据集,推动领域发展
组合图像检索(CIR)是一种重要范式,允许用户通过包含参考图像和修改文本的多模态查询检索目标图像。尽管当前研究进展显著,但主流方法仍依赖于仅覆盖有限显著变化的简单修改文本,导致两大实际应用中的关键问题:实体覆盖不足与语句-实体错位。为解决这些问题并使CIR更贴近真实场景,我们构建了两个指令丰富的多修改数据集:M-FashionIQ 和 M-CIRR。同时,提出TEMA——首个专为多修改设计的CIR框架,同时兼容简单修改。在四个基准数据集上的大量实验表明,TEMA在原始及多修改场景中均表现优越,且在检索精度与计算效率间保持最佳平衡。代码与数据集已公开于 https://github.com/lee-zixu/ACL26-TEMA/。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image and modification text. Although research on CIR has made significant progress, prevailing setups still rely simple modification texts that typically cover only a limited range of salient changes, which induces two limitations highly relevant to practical applications, namely Insufficient Entity Coverage and Clause-Entity Misalignment. In order to address these issues and bring CIR closer to real-world use, we construct two instruction-rich multi-modification datasets, M-FashionIQ and M-CIRR. In addition, we propose TEMA, the Text-oriented Entity Mapping Architecture, which is the first CIR framework designed for multi-modification while also accommodating simple modifications. Extensive experiments on four benchmark datasets demonstrate that TEMA's superiority in both original and multi-modification scenarios, while maintaining an optimal balance between retrieval accuracy and computational efficiency. Our codes and constructed multi-modification dataset (M-FashionIQ and M-CIRR) are available at https://github.com/lee-zixu/ACL26-TEMA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。