arXiv:2501.15842cs.LGcs.AI2025-01

对比三种先进模型在不同数据集间的泛化能力,发现简单模型反而更鲁棒。

Beyond In-Distribution Performance: A Cross-Dataset Study of Trajectory Prediction Robustness

  • 用不同数据集交叉训练测试,考察模型外分布泛化能力。
  • 小模型因强先验在小数据训练大模型测试时表现最佳。
  • 大模型在大数据训练小数据测试时全都不行,提示评估需谨慎。

我们研究了三种在分布内性能相当但结构不同的前沿轨迹预测模型的分布外泛化能力。通过在Argoverse 2(A2)上训练并在Waymo Open Motion(WO)上测试,反之亦然,考察归纳偏置、训练数据量和数据增强策略的影响。结果发现:当在较小的A2数据集上训练、在较大的WO上测试时,参数最少且归纳偏置最强的模型展现出最优的分布外泛化能力;而在反向设置下(在更大的WO上训练,在更小的A2上测试),所有模型泛化性能均差,即便最强归纳偏置的模型也仅表现相对较好。我们讨论这一反直觉现象的可能原因,并对轨迹预测模型的设计与评测基准提出建议。

原文摘要 · Abstract (English)

We study the Out-of-Distribution (OoD) generalization ability of three SotA trajectory prediction models with comparable In-Distribution (ID) performance but different model designs. We investigate the influence of inductive bias, size of training data and data augmentation strategy by training the models on Argoverse 2 (A2) and testing on Waymo Open Motion (WO) and vice versa. We find that the smallest model with highest inductive bias exhibits the best OoD generalization across different augmentation strategies when trained on the smaller A2 dataset and tested on the large WO dataset. In the converse setting, training all models on the larger WO dataset and testing on the smaller A2 dataset, we find that all models generalize poorly, even though the model with the highest inductive bias still exhibits the best generalization ability. We discuss possible reasons for this surprising finding and draw conclusions about the design and test of trajectory prediction models and benchmarks.

轨迹预测泛化能力数据集对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。