arXiv:2410.08903cs.RO2024-10被引 7

为自动驾驶系统评估构建动态人类基准,考虑时空差异提升公平性

Dynamic Benchmarks: Spatial and Temporal Alignment for ADS Performance Evaluation

  • 基于事故与行驶里程数据,动态调整时空分布差异
  • 旧基准误差最高达47%,新方法更真实反映性能表现
  • 适合自动驾驶评测、交通政策制定者参考

目前美国部分城市已部署无驾驶员的SAE L4+级自动驾驶出行服务。这类系统运营区域和时间可能与人类驾驶分布存在偏差。现有评估基准仅做县级地理匹配,忽略时空差异。本文提出新方法,结合警方事故数据、人类车公里(VMT)数据及超过2000万英里Waymo无人乘员运营数据,在三个美国县区构建动态人类基准。空间调整显示:旧基准在旧金山偏差10%至47%,马里科帕县12%至20%,洛杉矶县-7%至34%;圣何塞时间调整后,事故率比原值低2%至高16%。结果表明,需校正时空混杂因素,才能实现更公平的自动驾驶系统评估。

原文摘要 · Abstract (English)

Deployed SAE level 4+ Automated Driving Systems (ADS) without a human driver are currently operational ride-hailing fleets on surface streets in the United States. This current use case and future applications of this technology will determine where and when the fleets operate, potentially resulting in a divergence from the distribution of driving of some human benchmark population within a given locality. Existing benchmarks for evaluating ADS performance have only done county-level geographical matching of the ADS and benchmark driving exposure in crash rates. This study presents a novel methodology for constructing dynamic human benchmarks that adjust for spatial and temporal variations in driving distribution between an ADS and the overall human driven fleet. Dynamic benchmarks were generated using human police-reported crash data, human vehicle miles traveled (VMT) data, and over 20 million miles of Waymo's rider-only (RO) operational data accumulated across three US counties. The spatial adjustment revealed significant differences across various severity levels in adjusted crash rates compared to unadjusted benchmarks with these differences ranging from 10% to 47% higher in San Francisco, 12% to 20% higher in Maricopa, and 7% lower to 34% higher in Los Angeles counties. The time-of-day adjustment in San Francisco, limited to this region due to data availability, resulted in adjusted crash rates 2% lower to 16% higher than unadjusted rates, depending on severity level. The findings underscore the importance of adjusting for spatial and temporal confounders in benchmarking analysis, which ultimately contributes to a more equitable benchmark for ADS performance evaluations.

自动驾驶性能评估动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。