arXiv:2605.06012cs.CVcs.AI2026-05

提出细粒度图文车辆检索模型,提升文本描述查车的准确率

T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval

论文配图:T2I-VeRW: Part-level Fine-grained Perception for Text-to-Image Vehicle Retrieval
图 1 · 摘自论文原文
  • 按部件级别对图文进行局部对齐,引入可学习部件查询标记
  • 在T2I-VeRW数据集上达55.2% Rank-1准确率,领先现有方法
  • 适合需要高精度文本查车的应用场景,如交通监控与事故追查

车辆再识别(Re-ID)旨在从非重叠摄像头拍摄的图像中检索最相似的目标图像。将车辆Re-ID从仅图像查询扩展到文本查询,可在仅有目击者描述时实现真实场景下的检索。本文提出PFCVR模型,一种面向文本到图像车辆再识别的部件级细粒度跨模态检索方法。PFCVR在部件级别构建图文配对,并引入可学习的部件查询标记,在聚合部件特异性与全句上下文后与视觉部件特征对齐。在此显式局部对齐基础上,双向掩码恢复模块使各模态在对方引导下重建被遮蔽内容,隐式地将局部对应关系扩展为全局特征对齐。此外,我们构建了新的大规模数据集T2I-VeRW,包含14,668张图像、1,796个车辆身份,带有细粒度部件标注。在T2I-VeRI数据集上的实验表明,PFCVR达到29.2% Rank-1准确率,较最优竞争方法提升3.7个百分点。在新提出的T2I-VeRW基准上,PFCVR达到55.2% Rank-1准确率,优于一系列近期先进方法。

原文摘要 · Abstract (English)

Vehicle Re-identification (Re-ID) aims to retrieve the most similar image to a given query from images captured by non-overlapping cameras. Extending vehicle Re-ID from image-only queries to text-based queries enables retrieval in real-world scenarios where only a witness description of the target vehicle is available. In this paper, we propose PFCVR, a Part-level Fine-grained Cross-modal Vehicle Retrieval model for text-to-image vehicle re-identification. PFCVR constructs locally paired images and texts at the part level and introduces learnable part-query tokens that aggregate both part-specific and full-sentence context before aligning with visual part features. On top of this explicit local alignment, a bi-directional mask recovery module lets each modality reconstruct its masked content under the guidance of the other, implicitly bridging local correspondences into global feature alignment. Furthermore, we construct a new large-scale dataset called T2I-VeRW, which contains 14,668 images covering 1,796 vehicle identities with fine-grained part-level annotations. Experimental results on the T2I-VeRI dataset show that PFCVR achieves 29.2\% Rank-1 accuracy, improving over the best competing method by +3.7\% percentage points. On the newly proposed T2I-VeRW benchmark, PFCVR achieves 55.2\% Rank-1 accuracy, outperforming a comprehensive set of recent state-of-the-art methods. Source code will be released on https://github.com/Event-AHU/Neuromorphic_ReID

图文检索车辆识别细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。