arXiv:2411.18111cs.CV2024-11中稿 · ICASSP 2026被引 26

用大模型生成语义标记,提升跨摄像头行人识别准确率

When Large Vision-Language Models Meet Person Re-Identification

  • 用指令引导大模型生成包含关键外观信息的语义标记
  • 通过交互模块强化语义与视觉特征,提升身份表征能力
  • 无需额外图文标注,在多个数据集上表现优异

大型视觉语言模型(LVLM)在跨模态理解与推理任务中表现卓越。近年来,行人重识别(ReID)也开始探索跨模态语义以提高身份识别准确率。然而,如何有效利用LVLM进行ReID仍是开放挑战。尽管LVLM基于生成范式预测下一个输出词,但ReID需提取判别性身份特征以跨摄像头匹配行人。本文提出LVLM-ReID,一种新型框架,融合LVLM优势促进ReID。具体地,我们使用指令引导LVLM生成一个封装关键外观语义的语义标记,并通过语义引导交互(SGI)模块实现该标记与视觉标记间的双向交互,最终增强的语义标记作为行人身份表示。该框架将LVLM的语义理解与生成能力融入端到端ReID训练,使模型在训练与推理中均能捕捉丰富语义线索。实验表明,LVLM-ReID在多个基准上取得有竞争力的结果,且无需额外图像-文本标注,展示了LVLM生成语义在推进行人重识别中的潜力。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) that incorporate visual models and large language models have achieved impressive results across cross-modal understanding and reasoning tasks. In recent years, person re-identification (ReID) has also started to explore cross-modal semantics to improve the accuracy of identity recognition. However, effectively utilizing LVLMs for ReID remains an open challenge. While LVLMs operate under a generative paradigm by predicting the next output word, ReID requires the extraction of discriminative identity features to match pedestrians across cameras. In this paper, we propose LVLM-ReID, a novel framework that harnesses the strengths of LVLMs to promote ReID. Specifically, we employ instructions to guide the LVLM in generating one semantic token that encapsulates key appearance semantics from the person image. This token is further refined through our Semantic-Guided Interaction (SGI) module, establishing a reciprocal interaction between the semantic token and visual tokens. Ultimately, the reinforced semantic token serves as the representation of pedestrian identity. Our framework integrates the semantic understanding and generation capabilities of LVLM into end-to-end ReID training, allowing LVLM to capture rich semantic cues during both training and inference. LVLM-ReID achieves competitive results on multiple benchmarks without additional image-text annotations, demonstrating the potential of LVLM-generated semantics to advance person ReID.

行人重识别大模型跨模态语义生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。