arXiv:2609.06368cs.RO2026-09

构建闭环评估框架,量化协同预警对自动驾驶安全性提升效果

LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings

论文配图:LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings
图 1 · 摘自论文原文
  • 分离预警、危险、终止与恢复阶段,实现精准因果评估
  • 模型在预警下协作得分从34.6升至75.5,显著提升安全表现
  • 适合自动驾驶安全评估与视觉语言模型测试的研究者

我们提出LANTERN,一个用于时序定位协同预警的闭环评估基准。该基准将预警触发、危险发生、预警结束和灾后恢复阶段解耦,并在匹配的有预警与无预警场景下执行,以独立测量预警的贡献,避免车载视觉信息干扰。基准涵盖六类关键安全场景,包含3,272段序列共236,309帧训练数据,以及120组匹配的闭环评测路线。每条危险路线均在有预警与无预警条件下评估,其无危险对照组则惩罚非必要急刹。我们进一步提出协作统一评分(CUS),一种兼顾路线推进、预判、避让与恢复的安全门限指标。微调代表性视觉语言模型后,其在无预警下的CUS为34.6,启用预警后达75.5,验证了协同预警的价值与配对协议的区分能力。所有资源将公开可用。

原文摘要 · Abstract (English)

We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings. LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and evaluates each physical event under matched warning and no-warning executions so that the warning's contribution is measured in isolation rather than confounded with onboard vision. The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,309 frames for training, together with 120 matched route pairs for closed-loop evaluation. Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard control penalizes unconditional braking. We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewards route progress, anticipation, clearance, and recovery. Fine-tuning a representative VLM driving model raises CUS from 34.6 without warnings to 75.5 with them, demonstrating both the value of cooperative warnings and the discriminative power of the paired protocol. All resources will be made publicly available.

自动驾驶协同预警评估基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。