提出新任务Text-RGBT行人检索,融合可见光与热成像提升复杂环境识别能力。
Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark
- 分离视觉模态间关系,用双路径网络对齐文本与多模态图像特征
- 在4723对图像上实现SOTA性能,跨光照条件匹配准确率显著提升
- 构建高质量数据集RGBT-PEDES,含7987条细粒度文本描述
传统文本-图像行人检索易受光照变化影响,因可见光传感器成像局限。近年来,跨模态信息融合成为增强检索鲁棒性的有效策略。通过整合可见光与热成像的互补信息,可在复杂现实场景中实现更稳定的行人识别与匹配。受此启发,我们提出新任务:文本-可见光/热成像行人检索(Text-RGBT Person Retrieval),结合可见光与热成像的互补线索,提升挑战环境下行人检索能力。该任务的核心挑战在于对齐文本与多模态视觉特征,但可见光与热成像固有的异质性可能干扰视觉与语言间的对齐。为此,我们提出解耦式跨模态对齐网络(DCAlign),充分挖掘模态特异性与模态协同性视觉特征与文本的关系。为推动该领域发展,我们构建了高质量数据集RGBT-PEDES,包含1,822个身份(不同年龄与性别)、4,723对校准的RGB与热成像图像,覆盖昼夜多样场景,包含遮挡、弱对齐、恶劣光照等挑战。同时,为所有图像对精细标注7,987条细粒度文本描述。在RGBT-PEDES上的大量实验表明,所提方法优于现有文本-图像行人检索方法。
原文摘要 · Abstract (English)
The performance of traditional text-image person retrieval task is easily affected by lighting variations due to imaging limitations of visible spectrum sensors. In recent years, cross-modal information fusion has emerged as an effective strategy to enhance retrieval robustness. By integrating complementary information from different spectral modalities, it becomes possible to achieve more stable person recognition and matching under complex real-world conditions. Motivated by this, we introduce a novel task: Text-RGBT Person Retrieval, which incorporates cross-spectrum information fusion by combining the complementary cues from visible and thermal modalities for robust person retrieval in challenging environments. The key challenge of Text-RGBT person retrieval lies in aligning text with multi-modal visual features. However, the inherent heterogeneity between visible and thermal modalities may interfere with the alignment between vision and language. To handle this problem, we propose a Decoupled Cross-modal Alignment network (DCAlign), which sufficiently mines the relationships between modality-specific and modality-collaborative visual with the text, for Text-RGBT person retrieval. To promote the research and development of this field, we create a high-quality Text-RGBT person retrieval dataset, RGBT-PEDES. RGBT-PEDES contains 1,822 identities from different age groups and genders with 4,723 pairs of calibrated RGB and T images, and covers high-diverse scenes from both daytime and nighttime with a various of challenges such as occlusion, weak alignment and adverse lighting conditions. Additionally, we carefully annotate 7,987 fine-grained textual descriptions for all RGBT person image pairs. Extensive experiments on RGBT-PEDES demonstrate that our method outperforms existing text-image person retrieval methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。