提出首个跨模态行人重识别基准,验证多模态大模型潜力
Find Them All: Unveiling MLLMs for Versatile Person Re-identification
- 构建覆盖10类任务的多模态行人重识别基准VP-ReID
- 多模态大模型在多数任务上表现优异,但对热成像数据效果不佳
- 适合研究跨模态视觉与大模型融合的学者参考
行人重识别(ReID)旨在从图库中检索目标人物图像,广泛应用于医疗康复和公共安全。传统ReID模型多为单模态,跨异构数据模态时泛化能力有限。近年来,多模态大语言模型(MLLMs)展现出解决该问题的潜力,但现有方法仅将其用作特征提取器或图像描述生成器,未充分挖掘其在行人重识别中的能力。为此,我们提出了首个面向通用行人重识别的基准——VP-ReID,包含257,310个跨模态查询与图库图像,涵盖10种不同任务。同时设计了两种面向任务的评估方案。大量实验表明,MLLM在多种任务中表现出出色的泛化性、有效性与可解释性,但在热成像和红外等少数模态上仍存在局限。我们希望该基准能推动更鲁棒、通用的跨模态基础模型在行人重识别领域的研究。
原文摘要 · Abstract (English)
Person re-identification (ReID) aims to retrieve images of a target person from the gallery set, with wide applications in medical rehabilitation and public security. However, traditional person ReID models are typically uni-modal, resulting in limited generalizability across heterogeneous data modalities. Recently, the emergence of multi-modal large language models (MLLMs) has shown a promising avenue for addressing this issue. Despite this potential, existing methods merely regard MLLMs as feature extractors or caption generators, leaving their capabilities in person ReID tasks largely unexplored. To bridge this gap, we introduce a novel benchmark for \underline{\textbf{V}}ersatile \underline{\textbf{P}}erson \underline{\textbf{Re}}-\underline{\textbf{ID}}entification, termed VP-ReID. The benchmark includes 257,310 multi-modal queries and gallery images, covering ten diverse person ReID tasks. In addition, we propose two task-oriented evaluation schemes for MLLM-based person ReID. Extensive experiments demonstrate the impressive versatility, effectiveness, and interpretability of MLLMs in various person ReID tasks. Nevertheless, they also have limitations in handling a few modalities, particularly thermal and infrared data. We hope that VP-ReID can facilitate the community in developing more robust and generalizable cross-modal foundation models for person ReID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。