从传统特征到大模型,人重识别实现视觉与语义融合
Evolution of ReID: From Early Methods to LLM Integration
- 用GPT-4o生成动态身份提示,增强图像与文本对齐
- 在复杂场景下,语义描述使识别准确率显著提升
- 适合关注跨模态融合与未来方向的研究者
人重识别(ReID)已从手工特征方法演进至深度学习,并进一步发展为融合大语言模型(LLMs)的新范式。早期方法难以应对光照、姿态和视角变化,而深度学习通过学习鲁棒视觉特征缓解了这一问题。近期,LLMs使系统能利用自然语言中的语义与上下文信息提升匹配能力。本综述首次全面梳理了融合LLM的ReID方法,将文本描述作为先验信息以改进视觉匹配。关键贡献在于使用GPT-4o生成动态、身份特定的提示,显著增强视觉-语言对齐。实验表明,该方法在复杂或模糊情况下有效提升准确率。为支持后续研究,我们公开了标准ReID数据集上由GPT-4o生成的大规模描述数据集。本工作连接计算机视觉与自然语言处理,提供统一视角,并提出未来方向:更优提示设计、跨模态迁移学习及现实环境适应性。
原文摘要 · Abstract (English)
Person re-identification (ReID) has evolved from handcrafted feature-based methods to deep learning approaches and, more recently, to models incorporating large language models (LLMs). Early methods struggled with variations in lighting, pose, and viewpoint, but deep learning addressed these issues by learning robust visual features. Building on this, LLMs now enable ReID systems to integrate semantic and contextual information through natural language. This survey traces that full evolution and offers one of the first comprehensive reviews of ReID approaches that leverage LLMs, where textual descriptions are used as privileged information to improve visual matching. A key contribution is the use of dynamic, identity-specific prompts generated by GPT-4o, which enhance the alignment between images and text in vision-language ReID systems. Experimental results show that these descriptions improve accuracy, especially in complex or ambiguous cases. To support further research, we release a large set of GPT-4o-generated descriptions for standard ReID datasets. By bridging computer vision and natural language processing, this survey offers a unified perspective on the field's development and outlines key future directions such as better prompt design, cross-modal transfer learning, and real-world adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。