用文本描述引导跨模态对齐,提升可见光与红外行人重识别性能。
Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-Identification
- 通过多样语义描述引导跨模态特征对齐,增强身份一致性。
- 在SYSU-MM01、CUHK-CCXML、MLF-2400数据集上显著超越现有方法。
- 适合关注跨模态行人重识别与语义引导特征学习的研究者。
可见光-红外行人重识别(VI-ReID)因可见光与红外图像间存在巨大模态差异而面临挑战,导致特征难以对齐至统一空间。此外,光照、色彩对比等风格噪声会降低特征的身份判别力与模态不变性。为此,提出一种新颖的多样性语义引导特征对齐与解耦网络(DSFAD),将不同模态中与身份相关的特征对齐到文本嵌入空间,并解耦每模态内与身份无关的特征。设计多样性语义引导特征对齐(DSFA)模块,生成具有多样化句式结构的行人描述,以指导视觉特征的跨模态对齐。为过滤风格信息,提出语义边界引导特征解耦(SMFD)模块,将视觉特征分解为行人相关与风格相关成分,并约束前者与文本嵌入的相似度至少比后者高出一个边界值。为防止解耦过程丢失行人语义,设计语义一致性引导特征复原(SCFR)模块,从风格相关特征中挖掘有用信息并恢复至行人相关特征,同时确保复原后特征与文本嵌入的相似度与解耦前一致。在三个VI-ReID数据集(SYSU-MM01、CUHK-CCXML、MLF-2400)上的大量实验验证了所提方法的优越性。
原文摘要 · Abstract (English)
Visible-Infrared Person Re-Identification (VI-ReID) is a challenging task due to the large modality discrepancy between visible and infrared images, which complicates the alignment of their features into a suitable common space. Moreover, style noise, such as illumination and color contrast, reduces the identity discriminability and modality invariance of features. To address these challenges, we propose a novel Diverse Semantics-guided Feature Alignment and Decoupling (DSFAD) network to align identity-relevant features from different modalities into a textual embedding space and disentangle identity-irrelevant features within each modality. Specifically, we develop a Diverse Semantics-guided Feature Alignment (DSFA) module, which generates pedestrian descriptions with diverse sentence structures to guide the cross-modality alignment of visual features. Furthermore, to filter out style information, we propose a Semantic Margin-guided Feature Decoupling (SMFD) module, which decomposes visual features into pedestrian-related and style-related components, and then constrains the similarity between the former and the textual embeddings to be at least a margin higher than that between the latter and the textual embeddings. Additionally, to prevent the loss of pedestrian semantics during feature decoupling, we design a Semantic Consistency-guided Feature Restitution (SCFR) module, which further excavates useful information for identification from the style-related features and restores it back into the pedestrian-related features, and then constrains the similarity between the features after restitution and the textual embeddings to be consistent with that between the features before decoupling and the textual embeddings. Extensive experiments on three VI-ReID datasets demonstrate the superiority of our DSFAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。