arXiv:2509.14801cs.LG2025-09被引 6

STEP框架统一评测轨迹预测模型,揭示现有方法的缺陷与漏洞。

STEP: Structured Training and Evaluation Platform for benchmarking trajectory prediction models

  • 提供统一接口支持多数据集与多种模型,确保训练评估一致
  • 发现主流测试流程存在缺陷,联合建模能显著提升交互预测性能
  • 验证顶尖模型对分布偏移和对抗攻击均敏感,适合研究者深入分析

轨迹预测在自动驾驶路径规划中至关重要,但标准化评估方法仍不完善。尽管已有工作试图统一数据集格式与模型接口以促进比较,现有框架仍难以支持异构交通场景、联合预测模型或用户文档。本文提出STEP——一个新基准框架,通过统一多数据集接口、强制一致训练与评估条件,并支持广泛预测模型来解决上述问题。实验揭示:1)广泛使用的测试流程存在局限;2)联合建模代理对提升交互预测效果至关重要;3)当前最先进模型对分布偏移和对抗性代理攻击均表现出脆弱性。我们希望通过STEP推动研究从‘排行榜’竞争转向对复杂多智能体环境下模型行为与泛化能力的深层理解。

原文摘要 · Abstract (English)

While trajectory prediction plays a critical role in enabling safe and effective path-planning in automated vehicles, standardized practices for evaluating such models remain underdeveloped. Recent efforts have aimed to unify dataset formats and model interfaces for easier comparisons, yet existing frameworks often fall short in supporting heterogeneous traffic scenarios, joint prediction models, or user documentation. In this work, we introduce STEP -- a new benchmarking framework that addresses these limitations by providing a unified interface for multiple datasets, enforcing consistent training and evaluation conditions, and supporting a wide range of prediction models. We demonstrate the capabilities of STEP in a number of experiments which reveal 1) the limitations of widely-used testing procedures, 2) the importance of joint modeling of agents for better predictions of interactions, and 3) the vulnerability of current state-of-the-art models against both distribution shifts and targeted attacks by adversarial agents. With STEP, we aim to shift the focus from the ``leaderboard'' approach to deeper insights about model behavior and generalization in complex multi-agent settings.

轨迹预测多智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。