arXiv:2506.09385cs.CV2025-06NeurIPS被引 9

一个模型搞定所有模态的人体重识别,支持任意组合查询。

ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model

  • 用统一编码+多专家路由实现任意模态组合的融合与对齐。
  • 在包含1000人、5种模态的ORBench数据集上达到最佳性能。
  • 适合需要跨模态检索的实际场景,如安防与智能监控。

现实场景中,人体重识别(ReID)需通过描述性查询识别目标人物,无论查询是单模态还是多模态组合。但现有方法和数据集受限于有限模态,难以满足需求。为此,我们提出全新的全模态人体重识别(OM-ReID)问题,旨在实现不同多模态查询下的高效检索。为解决数据稀缺问题,我们构建了首个高质量多模态数据集ORBench,涵盖1000个唯一身份,包含五种模态:RGB、红外、彩色素描、素描和文本描述。该数据集在多样性方面具有显著优势,包括绘画视角和文本信息的丰富性,可作为后续研究的理想平台。此外,我们提出ReID5o,一种新颖的多模态学习框架,可在单一模型中实现任意模态组合的协同融合与跨模态对齐,采用统一编码与多专家路由机制。大量实验验证了ORBench的有效性与实用性,多种模型在该数据集上被评估比较,所提ReID5o模型表现最优。数据集与代码将公开于https://github.com/Zplusdragon/ReID5o_ORBench。

原文摘要 · Abstract (English)

In real-word scenarios, person re-identification (ReID) expects to identify a person-of-interest via the descriptive query, regardless of whether the query is a single modality or a combination of multiple modalities. However, existing methods and datasets remain constrained to limited modalities, failing to meet this requirement. Therefore, we investigate a new challenging problem called Omni Multi-modal Person Re-identification (OM-ReID), which aims to achieve effective retrieval with varying multi-modal queries. To address dataset scarcity, we construct ORBench, the first high-quality multi-modal dataset comprising 1,000 unique identities across five modalities: RGB, infrared, color pencil, sketch, and textual description. This dataset also has significant superiority in terms of diversity, such as the painting perspectives and textual information. It could serve as an ideal platform for follow-up investigations in OM-ReID. Moreover, we propose ReID5o, a novel multi-modal learning framework for person ReID. It enables synergistic fusion and cross-modal alignment of arbitrary modality combinations in a single model, with a unified encoding and multi-expert routing mechanism proposed. Extensive experiments verify the advancement and practicality of our ORBench. A wide range of possible models have been evaluated and compared on it, and our proposed ReID5o model gives the best performance. The dataset and code will be made publicly available at https://github.com/Zplusdragon/ReID5o_ORBench.

重识别多模态统一模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。