arXiv:2503.21595cs.CV2025-03被引 2

图文融合模型提升行人重识别与精准分割,应对遮挡等难题

PS-ReID: Advancing Person Re-Identification and Precise Segmentation with Multimodal Retrieval

  • 双路异构编码分离查询与目标角色,分别提取身份特征与场景信息
  • 在超20万张图的M2ReID数据集上,图文联合检索准确率显著超越单模态方法
  • 适合需要高精度定位与跨模态匹配的实际安防场景应用

行人重识别(ReID)在安全监控和刑侦中至关重要。传统基于图像的ReID方法常受遮挡和光照变化影响,而文本可提供互补信息,但多模态融合仍不充分。为此,我们提出PS-ReID,一种结合图像与文本输入的多模态模型,突破仅使用裁剪行人图像的局限,聚焦全场景设置,引入包含分割的多模态ReID任务,实现复杂条件下对目标个体的精确特征提取。模型采用双路异构编码架构,明确区分查询与目标角色:查询分支捕捉身份判别性线索,目标分支进行整体场景推理;同时引入令牌级ReID损失,监督身份感知令牌,使检索与分割耦合生成空间精准且身份一致的掩码。为系统评估,我们构建了当前规模最大的全场景多模态ReID数据集M2ReID,含超过20万张图像、4,894个身份,支持多模态查询与高质量分割掩码。实验表明,PS-ReID在ReID与分割任务上均显著优于单模态查询模型,在遮挡、低光照、背景杂乱等真实场景下表现优异,提供了鲁棒且灵活的行人检索与分割方案。所有代码、模型与数据集将公开。

原文摘要 · Abstract (English)

Person re-identification (ReID) plays a critical role in applications such as security surveillance and criminal investigations. Most traditional image-based ReID methods face challenges including occlusions and lighting changes, while text provides complementary information to mitigate these issues. However, the integration of both image and text modalities remains underexplored. To address this gap, we propose {\bf PS-ReID}, a multimodal model that combines image and text inputs to enhance ReID performance. In contrast to existing ReID methods limited by cropped pedestrian images, our PS-ReID focuses on full-scene settings and introduces a multimodal ReID task that incorporates segmentation, enabling precise feature extraction of the queried individual, even under challenging conditions such as occlusion. To this end, our model adopts a dual-path asymmetric encoding scheme that explicitly separates query and target roles: the query branch captures identity-discriminative cues, while the target branch performs holistic scene reasoning. Additionally, a token-level ReID loss supervises identity-aware tokens, coupling retrieval and segmentation to yield masks that are both spatially precise and identity-consistent. To facilitate systematic evaluation, we construct M2ReID, currently the largest full-scene multimodal ReID dataset, with over 200K images and 4,894 identities, featuring multimodal queries and high-quality segmentation masks. Experimental results demonstrate that PS-ReID significantly outperforms unimodal query-based models in both ReID and segmentation tasks. The model excels in challenging real-world scenarios such as occlusion, low lighting, and background clutter, offering a robust and flexible solution for person retrieval and segmentation. All code, models, and datasets will be publicly available.

行人重识别多模态精准分割图文检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。