arXiv:2603.28888cs.RO2026-03

用轻量视觉语言模型实现自动驾驶低延迟语义异常检测。

A Semantic Observer Layer for Autonomous Vehicles: Pre-Deployment Feasibility Study of VLMs for Low-Latency Anomaly Detection

  • 部署量化后的视觉语言模型,实时监控驾驶中的语义异常。
  • 推理延迟约500毫秒,较原始模型提速50倍,满足实时性要求。
  • 适合关注自动驾驶安全与边缘计算部署的研究者和工程师。

语义异常——即依赖上下文的危险场景,是传统像素级检测无法识别的关键安全风险。本文提出「语义观察层」:一个在自动驾驶控制回路旁以1-2赫兹运行的量化视觉语言模型(VLM),用于监测语义边缘情况,并在检测到时触发安全接管。采用Nvidia Cosmos-Reason1-7B模型,结合NVFP4量化与FlashAttention2,在相同硬件上实现约500毫秒的推理时间,相较未优化的FP16基线(无量化、标准PyTorch注意力)提速约50倍,满足观察器时序预算。我们在静态与视频条件下评估了准确率、延迟及量化行为,发现NF4量化导致召回率下降至10.6%,构成关键部署瓶颈,并通过危害分析将性能指标映射至安全目标。结果为该语义观察架构在具身智能自动驾驶平台上的预部署可行性提供了实证支持。

原文摘要 · Abstract (English)

Semantic anomalies-context-dependent hazards that pixel-level detectors cannot reason about-pose a critical safety risk in autonomous driving. We propose a \emph{semantic observer layer}: a quantized vision-language model (VLM) running at 1--2\,Hz alongside the primary AV control loop, monitoring for semantic edge cases, and triggering fail-safe handoffs when detected. Using Nvidia Cosmos-Reason1-7B with NVFP4 quantization and FlashAttention2, we achieve ~500 ms inference a ~50x speedup over the unoptimized FP16 baseline (no quantization, standard PyTorch attention) on the same hardware--satisfying the observer timing budget. We benchmark accuracy, latency, and quantization behavior in static and video conditions, identify NF4 recall collapse (10.6%) as a hard deployment constraint, and a hazard analysis mapping performance metrics to safety goals. The results establish a pre-deployment feasibility case for the semantic observer architecture on embodied-AI AV platforms.

自动驾驶语义检测量化模型安全系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。