arXiv:2409.00301cs.CV2024-09中稿 · the 27th IEEE Inte…被引 18

用视觉语言模型实现自动驾驶环境上下文的零样本识别

ContextVLM: Zero-Shot and Few-Shot Context Understanding for Autonomous Driving using Vision Language Models

  • 基于视觉语言模型,实现零样本和少样本环境上下文识别
  • 在160万数据对上准确率达95%以上,单次推理延迟仅10.5毫秒
  • 适合需要快速适应复杂路况的自动驾驶系统部署

近年来,自动驾驶技术不断发展以提升交通系统安全性。尽管部分自动驾驶车辆已在实际中部署,但全面应用仍需应对暴雨、大雪、低光照、施工区域及隧道内信号丢失等挑战。为此,自动驾驶车辆必须可靠识别运行环境的物理属性。本文将环境上下文识别定义为准确识别环境特征以应对相应情况的任务,共定义24种涵盖天气、光照、交通与道路状况的环境上下文。为支持该任务,我们构建了DrivingContexts数据集,包含超过160万条与自动驾驶相关的上下文-查询对。针对传统监督学习方法难以扩展至多样环境的问题,我们提出ContextVLM框架,利用视觉语言模型实现零样本与少样本上下文检测。该框架在自建数据集上达到95%以上准确率,且可在4GB Nvidia GeForce GTX 1050 Ti GPU上实时运行,单次查询延迟仅为10.5毫秒。

原文摘要 · Abstract (English)

In recent years, there has been a notable increase in the development of autonomous vehicle (AV) technologies aimed at improving safety in transportation systems. While AVs have been deployed in the real-world to some extent, a full-scale deployment requires AVs to robustly navigate through challenges like heavy rain, snow, low lighting, construction zones and GPS signal loss in tunnels. To be able to handle these specific challenges, an AV must reliably recognize the physical attributes of the environment in which it operates. In this paper, we define context recognition as the task of accurately identifying environmental attributes for an AV to appropriately deal with them. Specifically, we define 24 environmental contexts capturing a variety of weather, lighting, traffic and road conditions that an AV must be aware of. Motivated by the need to recognize environmental contexts, we create a context recognition dataset called DrivingContexts with more than 1.6 million context-query pairs relevant for an AV. Since traditional supervised computer vision approaches do not scale well to a variety of contexts, we propose a framework called ContextVLM that uses vision-language models to detect contexts using zero- and few-shot approaches. ContextVLM is capable of reliably detecting relevant driving contexts with an accuracy of more than 95% on our dataset, while running in real-time on a 4GB Nvidia GeForce GTX 1050 Ti GPU on an AV with a latency of 10.5 ms per query.

自动驾驶视觉语言模型上下文识别零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。