从单模到多模,统一梳理行人重识别技术演进
Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

- 系统梳理可见光-红外、文本-图像等跨模态重识别方法
- 提出基于Transformer的可见光-红外重识别基线框架
- 适合关注智能监控与多模态融合的科研人员阅读
行人重识别(ReID)是智能监控系统的关键组件,旨在跨不同摄像头网络匹配同一身份。传统方法主要依赖单模态可见光图像,易受光照不足和遮挡影响。为克服这些局限,领域正快速向跨模态与多模态范式发展。本综述系统回顾了可见光-红外(VI-ReID)、文本-图像(TI-ReID)、草图-图像(Sketch-ReID)以及新兴的非视距(NLOS)ReID等关键跨模态任务。同时探讨三波段与多模态融合ReID,分析异构传感器互补信息如何提升鲁棒性。除总结数据集、挑战与方法外,还提出一个基于Transformer的可见光-红外重识别基线框架,以有效捕捉模态不变特征。最后,基于当前研究格局,展望未来若干有前景的方向。
原文摘要 · Abstract (English)
Person re-identification (ReID) serves as a critical component in intelligent surveillance systems, aiming to match identities across disjoint camera networks. While traditional methods primarily rely on single-modal RGB imagery, they are often constrained by environmental challenges such as low illumination and occlusion. To overcome these limitations, the field is rapidly evolving toward cross-modal and multi-modal paradigms. This survey presents a comprehensive overview of this transition, systematically reviewing key cross-modal tasks including visible-infrared (VI-ReID), text-image (TI-ReID), sketch-based (Sketch-ReID), and the emerging Non-Line-of-Sight (NLOS) ReID, which extends perception beyond direct visibility. Furthermore, we examine tri-spectral and multi-modal fusion ReID, discussing how complementary information from diverse sensors enhances robustness. Beyond summarizing datasets, challenges, and methodologies, we propose a Transformer-based baseline framework for visible-infrared ReID, designed to effectively capture modality-invariant features. Finally, based on the current landscape, we outline several promising directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。