arXiv:2510.17685cs.CVcs.AI2025-10TPAMI被引 7

跨语言图文行人检索新框架,提升多语种下图文匹配精度

Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning

  • 双向隐式关系推理,自动建模图文局部关联
  • 多维全局对齐缓解跨模态差异,在多语言数据集上达到最佳性能
  • 首个多语言图文检索基准,适合多语种场景应用研究

文本到图像行人检索(TIPR)旨在通过文本描述定位目标行人,面临模态异构的挑战。现有方法采用跨模态全局或局部对齐策略,但全局方法忽略细粒度差异,局部方法需先验信息进行显式部件对齐。此外,当前方法以英语为中心,限制了多语言场景的应用。为此,我们首次提出多语言TIPR任务,构建多语言TIPR基准,利用大语言模型进行初步翻译,并结合领域知识进行优化。对应地,提出Bi-IRRA:双向隐式关系推理与对齐框架,通过双向掩码预测增强跨语言与跨模态的局部关系建模,并引入多维全局对齐模块缓解模态异构。该方法在所有多语言TIPR数据集上均取得新最优结果。数据与代码已公开于https://github.com/Flame-Chasers/Bi-IRRA。

原文摘要 · Abstract (English)

Text-to-image person retrieval (TIPR) aims to identify the target person using textual descriptions, facing challenge in modality heterogeneity. Prior works have attempted to address it by developing cross-modal global or local alignment strategies. However, global methods typically overlook fine-grained cross-modal differences, whereas local methods require prior information to explore explicit part alignments. Additionally, current methods are English-centric, restricting their application in multilingual contexts. To alleviate these issues, we pioneer a multilingual TIPR task by developing a multilingual TIPR benchmark, for which we leverage large language models for initial translations and refine them by integrating domain-specific knowledge. Correspondingly, we propose Bi-IRRA: a Bidirectional Implicit Relation Reasoning and Aligning framework to learn alignment across languages and modalities. Within Bi-IRRA, a bidirectional implicit relation reasoning module enables bidirectional prediction of masked image and text, implicitly enhancing the modeling of local relations across languages and modalities, a multi-dimensional global alignment module is integrated to bridge the modality heterogeneity. The proposed method achieves new state-of-the-art results on all multilingual TIPR datasets. Data and code are presented in https://github.com/Flame-Chasers/Bi-IRRA.

图文检索多语言跨模态行人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。