arXiv:2601.20391cs.IR2026-01

解决扩散模型生成图像时的幻觉问题,提升文本到图像检索准确率。

Eliminating Hallucination in Diffusion-Augmented Interactive Text-to-Image Retrieval

  • 通过语义一致性和扩散感知对比学习,让生成图像与文本对齐
  • 在五个基准上多轮检索命中率提升最高达7.37%
  • 适合需要高精度图文检索的应用场景

扩散增强的交互式文本到图像检索(DAI-TIR)通过扩散模型生成查询图像作为用户意图的额外视角,从而提升检索性能。然而,生成图像可能引入与原始文本冲突的幻觉视觉线索,导致性能下降。我们实证发现这些幻觉线索会显著损害DAI-TIR表现。为此,提出扩散感知多视图对比学习(DMCL),将DAI-TIR建模为查询意图与目标图像表示的联合优化。DMCL引入语义一致性与扩散感知对比目标,对齐文本与生成查询视图,同时抑制幻觉信号。该框架使编码器成为语义滤波器,将幻觉特征映射至零空间,增强对虚假线索的鲁棒性,更准确表达用户意图。注意力可视化与嵌入空间几何分析验证了此过滤行为。在五个标准基准上,DMCL在多轮Hits@10指标上持续提升,相较先前微调和零样本基线最高提升7.37%,表明其为DAI-TIR的通用且鲁棒的训练框架。

原文摘要 · Abstract (English)

Diffusion-Augmented Interactive Text-to-Image Retrieval (DAI-TIR) is a promising paradigm that improves retrieval performance by generating query images via diffusion models and using them as additional ``views'' of the user's intent. However, these generative views can be incorrect because diffusion generation may introduce hallucinated visual cues that conflict with the original query text. Indeed, we empirically demonstrate that these hallucinated cues can substantially degrade DAI-TIR performance. To address this, we propose Diffusion-aware Multi-view Contrastive Learning (DMCL), a hallucination-robust training framework that casts DAI-TIR as joint optimization over representations of query intent and the target image. DMCL introduces semantic-consistency and diffusion-aware contrastive objectives to align textual and diffusion-generated query views while suppressing hallucinated query signals. This yields an encoder that acts as a semantic filter, effectively mapping hallucinated cues into a null space, improving robustness to spurious cues and better representing the user's intent. Attention visualization and geometric embedding-space analyses corroborate this filtering behavior. Across five standard benchmarks, DMCL delivers consistent improvements in multi-round Hits@10, reaching as high as 7.37\% over prior fine-tuned and zero-shot baselines, which indicates it is a general and robust training framework for DAI-TIR.

图像检索扩散模型幻觉消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。