arXiv:2603.21886cs.IRcs.CV2026-03

改进图文检索的融合机制,让生成图像更精准匹配用户需求。

ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval

  • 用动态门控和专家混合机制,智能调节文本与生成图像的权重。
  • 在四个基准上提升检索效果,最高达3.49%的命中率增长。
  • 适合需要快速适配新查询的交互式图文检索系统使用。

近年来,交互式文本到图像检索(I-TIR)利用扩散模型弥合文本需求与待搜图像之间的模态差距,显著提升了检索效果。然而,现有方法通过简单的嵌入相加融合用户反馈的多模态信息,这种静态且无差别的融合方式会不加区分地引入扩散模型生成的噪声,导致高达55.62%样本性能下降。为此,本文提出ADaFuSE(语义感知专家的自适应扩散-文本融合),一种轻量级融合模型,可无缝接入现有框架而无需修改主干编码器。该模型采用双分支结构:其一为自适应门控分支,动态平衡不同模态的可靠性;其二为语义感知专家混合分支,捕捉细粒度跨模态特征。在四个标准I-TIR基准上的实验表明,ADaFuSE实现当前最优性能,在Hits@10上比DAR高出最多3.49%,参数仅增加5.29%,同时对噪声及长查询更具鲁棒性。结果表明,结合生成增强与合理融合,是一种简单且通用的替代微调的交互检索方案。

原文摘要 · Abstract (English)

Recent advances in interactive text-to-image retrieval (I-TIR) use diffusion models to bridge the modality gap between the textual information need and the images to be searched, resulting in increased effectiveness. However, existing frameworks fuse multi-modal views of user feedback by simple embedding addition. In this work, we show that this static and undifferentiated fusion indiscriminately incorporates generative noise produced by the diffusion model, leading to performance degradation for up to 55.62% samples. We further propose ADaFuSE (Adaptive Diffusion-Text Fusion with Semantic-aware Experts), a lightweight fusion model designed to align and calibrate multi-modal views for diffusion-augmented I-TIR, which can be plugged into existing frameworks without modifying the backbone encoder. Specifically, we introduce a dual-branch fusion mechanism that employs an adaptive gating branch to dynamically balance modality reliability, alongside a semantic-aware mixture-of-experts branch to capture fine-grained cross-modal nuances. Via thorough evaluation over four standard I-TIR benchmarks, ADaFuSE achieves state-of-the-art performance, surpassing DAR by up to 3.49% in Hits@10 with only a 5.29% parameter increase, while exhibiting stronger robustness to noisy and longer interactive queries. These results show that generative augmentation coupled with principled fusion provides a simple, generalizable alternative to fine-tuning for interactive retrieval.

图文检索扩散模型融合机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。