arXiv:2608.24094cs.RO2026-08

构建可调控应急车辆交互的仿真平台,用于评测自动驾驶安全性能。

SIREN-Bench: Behavior-Driven Generation and Evaluation of Emergency-Vehicle Interactions

论文配图:SIREN-Bench: Behavior-Driven Generation and Evaluation of Emergency-Vehicle Interactions
图 1 · 摘自论文原文
  • 用SUMO与CARLA协同仿真,按行为类型分控应急车与民用车
  • 7种交互模板覆盖3级紧急程度,支持传感器同步与标注
  • 发现检测与预测任务在不同行为下表现差异大,适合安全研究

应急车辆(EMVs)可通过使民用车辆刹车、变道或形成救援通道来重新组织周边交通。评估这类安全关键交互需对应急车辆特权与民用车辆响应实现行为级控制,并具备一致的感知数据与真实标签。现有数据集与仿真基准无法直接提供此组合。本文提出SIREN——一个面向应急车辆-民用车辆交互生成与评估的行为驱动型SUMO-CARLA协同仿真平台。SIREN将SUMO的网络级交通演化与行为逻辑,与CARLA的连续车辆控制及同步车载感知相结合;根据激活行为,交互由SUMO、CARLA或二者联合控制。我们构建了SIREN-Bench-v1,包含七种参数化交互模板,覆盖应急等级L1–L3和三类行为模式,支持同步传感器观测与仿真原生标注。通过三个代表性任务(3D目标检测、轨迹预测、视觉-语言风险理解)验证该基准。对九个轨迹预测器、四个基于LiDAR的检测器和五个视觉-语言模型的评估显示其行为依赖性失效模式:交通清空交互最难检测,优先交叉口通行最难预测,且无一学习模型平均优于恒定速度基线。视觉-语言模型在正常交通中表现显著优于近事故与碰撞事件。结果表明行为中心基准的价值,并确立SIREN为自动驾驶与交通安全部研究的可扩展数据生成与评估平台。

原文摘要 · Abstract (English)

Emergency vehicles (EMVs) can reorganize surrounding traffic as civilian vehicles brake, change lanes, or form rescue corridors in response to their passage. Evaluating these safety-critical interactions requires behavior-level control over both EMV privileges and civilian responses, together with consistent sensing and ground truth. Existing datasets and simulation benchmarks do not directly provide this combination. We present \textbf{SIREN}, a behavior-driven SUMO--CARLA co-simulation platform for generating EMV--civilian interactions. SIREN couples SUMO's network-level traffic evolution and behavior logic with CARLA's continuous vehicle control and synchronized onboard sensing; depending on the active behavior, the interaction is controlled by SUMO, CARLA, or jointly. We instantiate the platform as \textbf{SIREN-Bench-v1}, comprising seven parameterized interaction templates across emergency levels L1--L3 and three behavior families, with synchronized sensor observations and simulator-native annotations. We demonstrate the benchmark through three representative tasks: 3D object detection, trajectory prediction, and vision-language risk understanding. Evaluations of nine trajectory predictors, four LiDAR-based detectors, and five vision-language models reveal behavior-dependent failure modes. Traffic-clearance interactions are hardest for detection, privileged intersection traversal is hardest for prediction, and no learned predictor outperforms the constant-velocity reference on average. Vision-language models perform substantially better on normal traffic than on near-miss and collision events. These results demonstrate the value of behavior-centered benchmarking and establish SIREN as an extensible data-generation and evaluation platform for autonomous-driving and transportation safety research.

自动驾驶仿真平台应急车辆行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。