arXiv:2604.18376cs.CV2026-04被引 2

用大模型生成多视角文本,解决图文检索中的语义漂移问题。

Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation

论文配图:Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation
图 1 · 摘自论文原文
  • 通过双分支提示生成语义等价但表达多样的文本变体。
  • 无需训练即提升模型准确率,在三个数据集上达最优性能。
  • 适合关注跨模态对齐鲁棒性的研究者和应用开发者。

在文本到图像的人体检索任务中,自然语言表达的多样性与视觉语义的隐含性常导致语义漂移问题:语义相同但表述不同的文本在嵌入空间中出现显著特征差异,从而降低图文对齐的鲁棒性。本文提出一种基于大语言模型(LLM)的语义补偿框架(MVR),通过多视角语义重构与特征补偿提升跨模态表示一致性。核心方法包括:多视角重构(MVR):采用双分支提示策略,结合关键特征引导(通过特征相似性提取视觉关键组件)与多样性感知重写,生成语义等价且分布多样化的文本变体;文本特征鲁棒性增强:无需训练的潜在空间补偿机制,通过多视角特征均值池化与残差连接抑制噪声干扰,有效捕捉‘语义回声’;视觉语义补偿:视觉语言模型生成多角度图像描述,并通过共享文本重构进一步弥补视觉语义鸿沟。实验表明,该方法在不进行训练的情况下显著提升原始模型性能,并在三个文本到图像人体检索数据集上达到最先进水平。

原文摘要 · Abstract (English)

In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the problem of Expression Drift, where semantically equivalent texts exhibit significant feature discrepancies in the embedding space due to phrasing variations, thereby degrading the robustness of image-text alignment. This paper proposes a semantic compensation framework (MVR) driven by Large Language Models (LLMs), which enhances cross-modal representation consistency through multi-view semantic reformulation and feature compensation. The core methodology comprises three components: Multi-View Reformulation (MVR): A dual-branch prompting strategy combines key feature guidance (extracting visually critical components via feature similarity) and diversity-aware rewriting to generate semantically equivalent yet distributionally diverse textual variants; Textual Feature Robustness Enhancement: A training-free latent space compensation mechanism suppresses noise interference through multi-view feature mean-pooling and residual connections, effectively capturing "Semantic Echoes"; Visual Semantic Compensation: VLM generates multi-perspective image descriptions, which are further enhanced through shared text reformulation to address visual semantic gaps. Experiments demonstrate that our method can improve the accuracy of the original model well without training and performs SOTA on three text-to-image person retrieval datasets.

图文检索语义补偿大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。