通过人体轮廓信息增强可见光与红外图像的匹配效果
ShapeSpeak: Body Shape-Aware Textual Alignment for Visible-Infrared Person Re-Identification
- 用人体分割模型提取轮廓并转为结构化文本描述
- 在SYSU-MM01和RegDB上准确率显著提升
- 适合关注跨模态行人重识别的研究者
可见光-红外行人重识别(VIReID)旨在匹配可见光与红外行人图像,但模态差异和身份特征复杂性带来挑战。现有方法仅依赖身份标签监督,难以充分提取高层语义信息。近期视觉-语言预训练模型被引入,通过生成文本描述增强语义建模,但未显式建模人体轮廓特征,而该特征对跨模态匹配至关重要。为此,我们提出体形感知文本对齐(BSaTa)框架,显式建模并利用体形信息以提升性能。具体地,设计体形文本对齐(BSTA)模块,使用人体解析模型提取体形信息,并通过CLIP转换为结构化文本表示;设计文本-视觉一致性正则化器(TVCR),确保体形文本表示与视觉体形特征对齐。此外,引入体形感知表示学习(SRL)机制,结合多文本监督与分布一致性约束,引导视觉编码器学习模态不变且具有判别性的身份特征,增强模态不变性。实验结果表明,该方法在SYSU-MM01和RegDB数据集上均取得更优性能,验证了其有效性。
原文摘要 · Abstract (English)
Visible-Infrared Person Re-identification (VIReID) aims to match visible and infrared pedestrian images, but the modality differences and the complexity of identity features make it challenging. Existing methods rely solely on identity label supervision, which makes it difficult to fully extract high-level semantic information. Recently, vision-language pre-trained models have been introduced to VIReID, enhancing semantic information modeling by generating textual descriptions. However, such methods do not explicitly model body shape features, which are crucial for cross-modal matching. To address this, we propose an effective Body Shape-aware Textual Alignment (BSaTa) framework that explicitly models and utilizes body shape information to improve VIReID performance. Specifically, we design a Body Shape Textual Alignment (BSTA) module that extracts body shape information using a human parsing model and converts it into structured text representations via CLIP. We also design a Text-Visual Consistency Regularizer (TVCR) to ensure alignment between body shape textual representations and visual body shape features. Furthermore, we introduce a Shape-aware Representation Learning (SRL) mechanism that combines Multi-text Supervision and Distribution Consistency Constraints to guide the visual encoder to learn modality-invariant and discriminative identity features, thus enhancing modality invariance. Experimental results demonstrate that our method achieves superior performance on the SYSU-MM01 and RegDB datasets, validating its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。