用视觉词典实现可控的开放词汇图像检索,更懂用户意图且结果更丰富。
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
- 将查询和候选图都表示为开放词汇视觉词典,统一嵌入空间匹配
- 在DFMM-Compose上召回率提升3.2点,加自监督预训练再增2.3点
- 适合需要精准语义控制与多样化结果的图像编辑场景
组合图像检索(CIR)通过参考图像与自然语言编辑指令,检索出应用指定修改并保留其他相关视觉内容的图像。传统融合方法依赖监督三元组,易丢失细粒度特征;近期零样本方法通常对参考图生成描述并合并编辑指令,可能忽略隐含用户意图,导致结果重复。本文提出Pix2Key,将查询与候选图像均表示为开放词汇视觉词典,在统一嵌入空间中实现意图感知的约束匹配与多样性感知的重排序。其自监督预训练模块V-Dict-AE仅使用图像数据优化词典表示,增强细粒度属性理解,无需特定任务标注。在DFMM-Compose基准测试中,Pix2Key使Recall@10最高提升3.2点,引入V-Dict-AE后进一步提升2.3点,同时改善意图一致性并保持高列表多样性。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) uses a reference image plus a natural-language edit to retrieve images that apply the requested change while preserving other relevant visual content. Classic fusion pipelines typically rely on supervised triplets and can lose fine-grained cues, while recent zero-shot approaches often caption the reference image and merge the caption with the edit, which may miss implicit user intent and return repetitive results. We present Pix2Key, which represents both queries and candidates as open-vocabulary visual dictionaries, enabling intent-aware constraint matching and diversity-aware reranking in a unified embedding space. A self-supervised pretraining component, V-Dict-AE, further improves the dictionary representation using only images, strengthening fine-grained attribute understanding without CIR-specific supervision. On the DFMM-Compose benchmark, Pix2Key improves Recall@10 up to 3.2 points, and adding V-Dict-AE yields an additional 2.3-point gain while improving intent consistency and maintaining high list diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。