让AI通过对话精准识别跨摄像头行人,支持灵活交互与细节推理。
ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models
- 以文本主导的渐进式微调框架,分三阶段提升模型理解力
- 在10个主流数据集上均达顶尖性能,显著优于现有方法
- 适合需要开放对话交互与细粒度识别的应用场景
行人重识别(Re-ID)是计算机视觉中的关键任务,旨在跨非重叠摄像头视图识别同一人物。尽管先进视觉语言模型(VLMs)在逻辑推理和多任务泛化方面表现优异,但其在Re-ID任务中的应用仍受限,或难以基于身份相关特征准确匹配,或仅作为图像主导分支的辅助语义。本文提出新框架ChatReID,转向以文本为主导的检索范式,实现灵活且可交互的重识别。首先构建包含超过800万条指令的大规模指令数据集,促进模型微调;其次提出分层渐进式微调策略,通过三个阶段(从人物属性理解到细粒度图像检索,再到多模态任务推理)逐步赋予模型Re-ID能力。在10个主流基准上的大量实验表明,ChatReID在所有Re-ID任务中均达到当前最优性能。更多实验显示,该模型不仅能识别细粒度细节,还能将其整合进连贯的推理过程。
原文摘要 · Abstract (English)
Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task generalization, their applications in Re-ID tasks remain limited. They either struggle to perform accurate matching based on identity-relevant features or assist image-dominated branches as auxiliary semantics. In this paper, we propose a novel framework ChatReID, that shifts the focus towards a text-side-dominated retrieval paradigm, enabling flexible and interactive re-identification. To integrate the reasoning abilities of language models into Re-ID pipelines, We first present a large-scale instruction dataset, which contains more than 8 million prompts to promote the model fine-tuning. Next. we introduce a hierarchical progressive tuning strategy, which endows Re-ID ability through three stages of tuning, i.e., from person attribute understanding to fine-grained image retrieval and to multi-modal task reasoning. Extensive experiments across ten popular benchmarks demonstrate that ChatReID outperforms existing methods, achieving state-of-the-art performance in all Re-ID tasks. More experiments demonstrate that ChatReID not only has the ability to recognize fine-grained details but also to integrate them into a coherent reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。