arXiv:2607.16943cs.RO2026-07

构建多城市无人机数据集,标注交通风险,助力自动驾驶安全验证。

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

论文配图:SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections
图 1 · 摘自论文原文
  • 六城6个路口采集,覆盖不同道路结构与驾驶习惯
  • 提取32,682个高危交互事件,丰富边界测试场景
  • 含违规、遮挡等多维语义标签,适合安全算法评估

自动驾驶系统在信号交叉口的安全验证仍是部署瓶颈,因这类场景涉及密集异构交通、路权争议及长尾高危交互,对功能安全(SOTIF)构成严峻挑战。现有自然驾驶数据集普遍存在地理同质性、高危事件稀疏、缺乏语义风险标注等问题,限制了算法泛化性评估与针对性SOTIF验证。为此,本文提出SinD 2.0,一个大规模基于无人机的交叉口数据集,专用于跨域自动驾驶安全分析。主要贡献包括:(1) 跨域多样性:覆盖中国四座城市中的六个信号交叉口,捕捉各异的道路拓扑与区域驾驶特征;(2) 高密度风险交互:通过代理安全度量提取32,682个安全关键事件,显著提升边界测试场景密度;(3) 分层语义标注:融合高清地图与信号相位时间(SPaT)数据,提供包含交通违规、高风险交互、视觉遮挡、狭窄可行区域等多维度语义标签;(4) 全栈测试工具链:支持自动化场景提取、仅预测评估、开环回放、反应式闭环测试及真实感渲染。基准实验表明,SinD 2.0在城市间表现出显著域偏移,其语义风险子集能有效暴露自动驾驶算法性能短板。数据集、标注与工具链已开源于https://github.com/SOTIF-AVLab/SinD/tree/main。

原文摘要 · Abstract (English)

Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenarios involve dense heterogeneous traffic, contested right of way, and long-tail safety-critical interactions, posing significant challenges to the Safety of the Intended Functionality (SOTIF). Existing naturalistic driving datasets often suffer from geographical homogeneity, sparsity of safety-critical events, and lack of semantic risk annotations, which limit the evaluation of algorithmic generalizability and targeted SOTIF verification. To address these gaps, this paper introduces SinD 2.0, a large-scale drone-based intersection dataset dedicated to cross-domain ADS safety analysis. The main contributions of SinD 2.0 are: (1) Cross-domain diversity: It covers six signalized intersections across four Chinese cities, capturing distinct intersection topologies and regional driving behavior characteristics; (2) High-density risk interactions: A total of 32,682 safety-critical events are extracted via surrogate safety measures, significantly enriching the density of boundary test scenarios; (3) Hierarchical semantic annotations: Besides integration with high-definition (HD) maps and Signal Phase and Timing (SPaT) data, it provides multi-dimensional semantic labels including traffic violations, high-risk interactions, visual shielding, and narrow feasible areas; (4) Full-stack testing toolchain: It supports automated scenario extraction, prediction-only evaluation, open-loop replay, reactive closed-loop testing, and photorealistic rendering. Benchmark experiments demonstrate that SinD 2.0 exhibits significant domain shifts across cities, and the semantic risk subsets can effectively expose the performance limitations of ADS algorithms. The dataset, annotations, and testing toolchain are available at https://github.com/SOTIF-AVLab/SinD/tree/main.

自动驾驶安全验证无人机数据集语义标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。