跨语言图文行人检索新框架,提升多语种下图文匹配精度
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- 双向隐式关系推理,自动建模图文局部关联
- 多维全局对齐缓解跨模态差异,在多语言数据集上达到最佳性能
- 首个多语言图文检索基准,适合多语种场景应用研究
文本到图像行人检索(TIPR)旨在通过文本描述定位目标行人,面临模态异构的挑战。现有方法采用跨模态全局或局部对齐策略,但全局方法忽略细粒度差异,局部方法需先验信息进行显式部件对齐。此外,当前方法以英语为中心,限制了多语言场景的应用。为此,我们首次提出多语言TIPR任务,构建多语言TIPR基准,利用大语言模型进行初步翻译,并结合领域知识进行优化。对应地,提出Bi-IRRA:双向隐式关系推理与对齐框架,通过双向掩码预测增强跨语言与跨模态的局部关系建模,并引入多维全局对齐模块缓解模态异构。该方法在所有多语言TIPR数据集上均取得新最优结果。数据与代码已公开于https://github.com/Flame-Chasers/Bi-IRRA。
原文摘要 · Abstract (English)
Text-to-image person retrieval (TIPR) aims to identify the target person using textual descriptions, facing challenge in modality heterogeneity. Prior works have attempted to address it by developing cross-modal global or local alignment strategies. However, global methods typically overlook fine-grained cross-modal differences, whereas local methods require prior information to explore explicit part alignments. Additionally, current methods are English-centric, restricting their application in multilingual contexts. To alleviate these issues, we pioneer a multilingual TIPR task by developing a multilingual TIPR benchmark, for which we leverage large language models for initial translations and refine them by integrating domain-specific knowledge. Correspondingly, we propose Bi-IRRA: a Bidirectional Implicit Relation Reasoning and Aligning framework to learn alignment across languages and modalities. Within Bi-IRRA, a bidirectional implicit relation reasoning module enables bidirectional prediction of masked image and text, implicitly enhancing the modeling of local relations across languages and modalities, a multi-dimensional global alignment module is integrated to bridge the modality heterogeneity. The proposed method achieves new state-of-the-art results on all multilingual TIPR datasets. Data and code are presented in https://github.com/Flame-Chasers/Bi-IRRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。