arXiv:2606.09362cs.CVcs.LG2026-06

用视觉语言模型生成语义描述,实现无需训练的自动驾驶目标重识别

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

论文配图:Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study
图 1 · 摘自论文原文
  • 通过视觉语言模型提取物体语义属性进行匹配
  • 零样本下检索性能接近监督式CNN基线
  • 提升可解释性但存在视角不一致和细粒度区分不足问题

自动驾驶中的重识别通常作为视觉匹配问题处理,依赖学习到的外观嵌入在时间、帧或摄像头视角间关联车辆、行人和骑行者。然而,纯视觉表示易受视角、遮挡、光照和传感器域变化影响,限制其在复杂驾驶场景中的可解释性和鲁棒性。本文提出一种基于视觉语言模型(VLM)的零样本管道基准研究,通过生成检测到交通参与者的文本描述,评估这些描述是否能支持跨观测的身份匹配。该方法不再仅依赖低层视觉相似性,而是通过类别、颜色、形状、姿态、可见部分、空间上下文及显著视觉线索等结构化语义属性表征每个对象。实验表明,零样本语义描述可支持有效的对象重识别,在检索性能上与监督式CNN基线相当,同时通过显式身份线索提供更高可解释性。但实验也揭示了关键挑战:不同视角下属性不一致,且对视觉相似实例的细粒度区分能力有限。

原文摘要 · Abstract (English)

Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cyclists are associated across time, frames, or camera views using learned appearance embeddings, often complemented by motion, geometric, or multimodal cues. However, purely visual representations may be sensitive to viewpoint, occlusion, illumination, and sensor-domain variations, limiting their interpretability and robustness in complex driving scenes. We propose a baseline study of a zero-shot pipeline using Vision-Language Models (VLMs) to generate textual descriptions of detected traffic participants and evaluate whether these descriptions can support identity matching across observations. Instead of relying only on low-level visual similarity, the proposed formulation represents each object through structured semantic attributes, including category, color, shape, pose, visible parts, spatial context, and distinctive visual cues. This study provides an initial benchmark for language-based re-identification in autonomous-driving scenarios, discussing and evaluating the strengths and limitations of current VLMs for this task. Results demonstrate that zero-shot semantic descriptions can support effective object re-identification, achieving retrieval performance comparable to a supervised CNN baseline while offering greater interpretability through explicit identity cues. However, the experiments also reveal important challenges, including attribute inconsistency across viewpoints and limited fine-grained discrimination between visually similar instances.

重识别视觉语言模型自动驾驶零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。