arXiv:2502.16815cs.CV2025-02被引 4

用CLIP自动增强车辆细粒度语义,提升跨摄像头识别准确率

CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification

  • 基于CLIP提取通用语义,通过自适应模块精细增强特征
  • 在VeRi-776上达92.9% mAP、98.7% Rank-1,刷新纪录
  • 无需额外标注,适合无监督或弱监督场景下的车辆识别

车辆重识别(Re-ID)是智能交通系统中的关键任务,旨在跨不同监控摄像头检索和匹配同一辆车。现有方法常依赖额外标注信息来增强语义特征,存在局限性。本文提出一种基于CLIP的语义增强网络(CLIP-SENet),一个端到端框架,可自主提取并优化车辆语义属性,生成更鲁棒的语义特征表示。受大规模视觉-语言模型零样本能力启发,我们利用CLIP图像编码器提取通用语义信息,并设计自适应细粒度增强模块(AFEM),在细粒度层面动态优化该信息,获得更强语义表征。这些特征与常规外观特征融合,进一步区分不同车辆。在三个基准数据集上的全面评估表明,该方法有效:在VeRi-776上达到92.9% mAP、98.7% Rank-1;VehicleID上达90.4% Rank-1、98.7% Rank-5;VeRi-Wild上达89.1% mAP、97.9% Rank-1,均刷新当前最佳性能。

原文摘要 · Abstract (English)

Vehicle re-identification (Re-ID) is a crucial task in intelligent transportation systems (ITS), aimed at retrieving and matching the same vehicle across different surveillance cameras. Numerous studies have explored methods to enhance vehicle Re-ID by focusing on semantic enhancement. However, these methods often rely on additional annotated information to enable models to extract effective semantic features, which brings many limitations. In this work, we propose a CLIP-based Semantic Enhancement Network (CLIP-SENet), an end-to-end framework designed to autonomously extract and refine vehicle semantic attributes, facilitating the generation of more robust semantic feature representations. Inspired by zero-shot solutions for downstream tasks presented by large-scale vision-language models, we leverage the powerful cross-modal descriptive capabilities of the CLIP image encoder to initially extract general semantic information. Instead of using a text encoder for semantic alignment, we design an adaptive fine-grained enhancement module (AFEM) to adaptively enhance this general semantic information at a fine-grained level to obtain robust semantic feature representations. These features are then fused with common Re-ID appearance features to further refine the distinctions between vehicles. Our comprehensive evaluation on three benchmark datasets demonstrates the effectiveness of CLIP-SENet. Our approach achieves new state-of-the-art performance, with 92.9% mAP and 98.7% Rank-1 on VeRi-776 dataset, 90.4% Rank-1 and 98.7% Rank-5 on VehicleID dataset, and 89.1% mAP and 97.9% Rank-1 on the more challenging VeRi-Wild dataset.

车辆重识别CLIP语义增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。