用自动生成的轨迹预训练,让自动驾驶预测更鲁棒、更易推广。
PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting
- 用现成3D检测器生成伪标签轨迹,无需人工标注。
- 在少量标注数据下仍表现优异,跨场景泛化能力更强。
- 适合数据稀缺或需跨域部署的自动驾驶场景。
准确预测动态场景中智能体的运动是实现安全自动驾驶的关键。当前顶尖的运动预测模型依赖于人工标注或后处理的轨迹数据集,但构建这些数据集成本高、依赖人力、难以扩展且缺乏可复现性,还引入领域差距,限制了模型在不同环境中的泛化能力。本文提出PPT(Pretraining with Pseudo-labeled Trajectories),一种简单且可扩展的预训练框架,利用现成3D检测器和跟踪系统自动生成的未处理轨迹作为信号,进行模型预训练。与追求干净单标签标注的数据流水线不同,PPT主动利用现有轨迹中的多样性作为学习鲁棒表征的有用信号。通过在少量标注数据上进行微调,基于PPT预训练的模型在多个标准基准上表现强劲,尤其在低数据场景、跨域、端到端及多类别设置中优势显著。PPT易于实现,显著提升了运动预测的泛化性能。
原文摘要 · Abstract (English)
Accurately predicting how agents move in dynamic scenes is essential for safe autonomous driving. State-of-the-art motion forecasting models rely on datasets with manually annotated or post-processed trajectories. However, building these datasets is costly, generally manual, hard to scale, and lacks reproducibility. They also introduce domain gaps that limit generalization across environments. We introduce PPT (Pretraining with Pseudo-labeled Trajectories), a simple and scalable pretraining framework that uses unprocessed and diverse trajectories automatically generated from off-the-shelf 3D detectors and tracking. Unlike data annotation pipelines aiming for clean, single-label annotations, PPT is a pretraining framework embracing off-the-shelf trajectories as useful signals for learning robust representations. With optional finetuning on a small amount of labeled data, models pretrained with PPT achieve strong performance across standard benchmarks, particularly in low-data regimes, and in cross-domain, end-to-end, and multi-class settings. PPT is easy to implement and improves generalization in motion forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。