通过解耦概念表示,提升文本到图像的人体再识别准确率
DiCo: Disentangled Concept Representation for Text-to-image Person Re-identification
- 用槽位-概念块结构解耦颜色、纹理等属性
- 在三个数据集上达到顶尖性能,细粒度匹配更精准
- 适合需要可解释性的人像检索场景
文本到图像人体再识别(TIReID)旨在根据自由文本描述从大规模图库中检索人物图像。该任务因视觉外观与文本表达间存在显著模态差异,且需建模细粒度对应关系以区分相似属性(如服装颜色、纹理或风格)的人物而极具挑战。为此,我们提出DiCo(解耦概念表示)框架,实现分层解耦的跨模态对齐。DiCo引入共享槽位表示,每个槽位作为跨模态的局部锚点,并进一步分解为多个概念块,从而解耦互补属性(如颜色、纹理、形状),同时保持图像与文本间的稳定局部对应。在CUHK-PEDES、ICFG-PEDES和RSTPReid上的大量实验表明,该框架性能媲美最先进方法,且通过显式的槽位与块级表示提升了可解释性,支持更细粒度的检索结果。
原文摘要 · Abstract (English)
Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual expressions, as well as the need to model fine-grained correspondences that distinguish individuals with similar attributes such as clothing color, texture, or outfit style. To address these issues, we propose DiCo (Disentangled Concept Representation), a novel framework that achieves hierarchical and disentangled cross-modal alignment. DiCo introduces a shared slot-based representation, where each slot acts as a part-level anchor across modalities and is further decomposed into multiple concept blocks. This design enables the disentanglement of complementary attributes (\textit{e.g.}, color, texture, shape) while maintaining consistent part-level correspondence between image and text. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReid demonstrate that our framework achieves competitive performance with state-of-the-art methods, while also enhancing interpretability through explicit slot- and block-level representations for more fine-grained retrieval results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。