arXiv:2511.08978cs.MMcs.CV2025-11被引 10

将时空数据融入视觉语言模型,提升交通场景理解能力

Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding

  • 设计动态时空上下文提示机制,融合位置与时间信息
  • 在少样本条件下显著提升复杂交通场景理解准确率
  • 适合自动驾驶、智能导航等需多模态感知的场景

当前导航与网约车应用积累了大量带时空信息的图像。交通场景理解(TSU)的核心目标是全面描述交通环境,但传统方法常忽略时空数据,仅将其视为普通图像理解任务。为此,本文提出基于CLIP的时空增强模型ST-CLIP,引入一种双层时空感知多视角提示学习方法(SCAMP)。该方法包含:1)动态时空上下文表征模块,提取每张图像的时空特征向量;2)双层提示学习模块,将这些向量融入提示词嵌入,并结合低层视觉特征与高层语义特征,挖掘交通要素间的交互关系。实验在两个真实数据集上验证,该模型在少样本设置下表现优异,首次将时空信息系统性集成至视觉语言模型以支持TSU任务。

原文摘要 · Abstract (English)

Nowadays, navigation and ride-sharing apps have collected numerous images with spatio-temporal data. A core technology for analyzing such images, associated with spatiotemporal information, is Traffic Scene Understanding (TSU), which aims to provide a comprehensive description of the traffic scene. Unlike traditional spatio-temporal data analysis tasks, the dependence on both spatio-temporal and visual-textual data introduces distinct challenges to TSU task. However, recent research often treats TSU as a common image understanding task, ignoring the spatio-temporal information and overlooking the interrelations between different aspects of the traffic scene. To address these issues, we propose a novel SpatioTemporal Enhanced Model based on CILP (ST-CLIP) for TSU. Our model uses the classic vision-language model, CLIP, as the backbone, and designs a Spatio-temporal Context Aware Multiaspect Prompt (SCAMP) learning method to incorporate spatiotemporal information into TSU. The prompt learning method consists of two components: A dynamic spatio-temporal context representation module that extracts representation vectors of spatio-temporal data for each traffic scene image, and a bi-level ST-aware multi-aspect prompt learning module that integrates the ST-context representation vectors into word embeddings of prompts for the CLIP model. The second module also extracts low-level visual features and image-wise high-level semantic features to exploit interactive relations among different aspects of traffic scenes. To the best of our knowledge, this is the first attempt to integrate spatio-temporal information into visionlanguage models to facilitate TSU task. Experiments on two realworld datasets demonstrate superior performance in the complex scene understanding scenarios with a few-shot learning strategy.

交通理解视觉语言模型时空建模少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。