arXiv:2605.07649cs.CVcs.AI2026-05

用视觉语言模型零样本感知自动驾驶运行设计域,提升安全可审计性。

Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

论文配图:Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
图 1 · 摘自论文原文
  • 采用定义锚定的思维链提示,让模型零样本理解运行设计域
  • 在自建数据集与Mapillary Vistas上实现高精度零样本分类与检测
  • 提供可复用提示模板,适合自动驾驶安全验证场景

近年来,自主系统研究已发展到可面向特定领域落地应用的阶段。但大规模应用仍需满足安全法规要求,而这些法规常基于运行设计域(ODD)定义系统可工作的具体条件,尤其对自动驾驶系统而言,准确感知ODD要素是安全实施与审计的关键。视觉语言模型(VLMs)融合视觉识别与语言推理能力,无需任务特训数据即可工作,适用于动态变化的ODD感知。本文贡献:(i) 在自建数据集和Mapillary Vistas上,对四种VLM进行零样本ODD分类与检测的实证研究及失效分析;(ii) 对零样本优化策略进行消融实验,评估成本与性能;(iii) 提供一套可复用的提示模板及适配指导。结果表明,定义锚定的思维链提示结合角色分解效果最佳,其他方法可能导致召回率下降。整体成果为安全关键场景下的透明、高效ODD感知提供了可行路径。

原文摘要 · Abstract (English)

Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate research into practical, stakeholder-driven use cases across well-defined domains. However, for a wide-scale practical adoption of autonomous systems, adherence to safety regulations is crucial. Many regulations are influenced by the Operational Design Domain (ODD), which defines the specific conditions in which an autonomous agent can function. This is especially relevant for Automated Driving Systems (ADS), as a dependable perception of ODD elements is essential for safe implementation and auditing. Vision-language models (VLMs) integrate visual recognition and language reasoning, functioning without task-specific training data, which makes them suitable for adaptable ODD perception. To assess whether VLMs can function as zero-shot "ODD sensors" that adapt to evolving definitions, we contribute (i) an empirical study of zero-shot ODD classification and detection using four VLMs on a custom dataset and Mapillary Vistas, along with failure analyses; (ii) an ablation of zero-shot optimization strategies with a cost-performance overview; and (iii) a suite of reusable prompting templates with guidance for adaptation. Our findings indicate that definition-anchored chain-of-thought prompting with persona decomposition performs best, while other methods may result in reduced recall. Overall, our results pave the way for transparent and effective ODD-based perception in safety-critical applications.

自动驾驶视觉语言模型零样本安全感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。