arXiv:2510.10086cs.RO2025-10被引 1

提出新评估框架,更真实测试自动驾驶预测模型在复杂场景下的表现。

Beyond ADE and FDE: A Comprehensive Evaluation Framework for Safety-Critical Prediction in Multi-Agent Autonomous Driving Scenarios

  • 设计多场景测试框架,区分近距与远距车辆影响
  • 发现现有模型在高密度交互下易出错,传统指标无法暴露问题
  • 适合安全关键系统研发者和评测人员使用

当前自动驾驶预测模型的评估主要依赖平均位移误差(ADE)和最终位移误差(FDE)等简单指标。这些指标虽能提供基础性能评估,却难以捕捉模型在复杂、交互性强且关乎安全的驾驶场景中的真实表现。例如,现有基准未能区分邻近与远处车辆的影响,也未系统测试模型在不同多智能体交互下的鲁棒性。本文提出一种新型测试框架,评估模型在多样场景结构、地图上下文、车辆密度及空间分布下的表现。通过大量实证分析,量化了车辆距离对目标轨迹预测的影响,并识别出传统指标无法揭示的场景特异性失败案例。研究结果揭示了当前最先进预测模型的关键缺陷,强调了场景感知评估的重要性。该框架为安全驱动的预测验证奠定了基础,有助于发现潜在风险边界情况,推动可认证、鲁棒的自动驾驶预测系统发展。

原文摘要 · Abstract (English)

Current evaluation methods for autonomous driving prediction models rely heavily on simplistic metrics such as Average Displacement Error (ADE) and Final Displacement Error (FDE). While these metrics offer basic performance assessments, they fail to capture the nuanced behavior of prediction modules under complex, interactive, and safety-critical driving scenarios. For instance, existing benchmarks do not distinguish the influence of nearby versus distant agents, nor systematically test model robustness across varying multi-agent interactions. This paper addresses this critical gap by proposing a novel testing framework that evaluates prediction performance under diverse scene structures, saying, map context, agent density and spatial distribution. Through extensive empirical analysis, we quantify the differential impact of agent proximity on target trajectory prediction and identify scenario-specific failure cases that are not exposed by traditional metrics. Our findings highlight key vulnerabilities in current state-of-the-art prediction models and demonstrate the importance of scenario-aware evaluation. The proposed framework lays the groundwork for rigorous, safety-driven prediction validation, contributing significantly to the identification of failure-prone corner cases and the development of robust, certifiable prediction systems for autonomous vehicles.

自动驾驶预测评估多智能体安全关键

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。