解耦表征与重建,提升多智能体轨迹预测的场景一致性
DECAMP: Towards Scene-Consistent Multi-Agent Motion Prediction with Disentangled Context-Aware Pre-Training
- 分离行为模式学习与特征重建,增强可解释性
- 在Argoverse 2上实现更优的多智能体预测性能
- 适合自动驾驶场景下的轨迹预测研究者
轨迹预测是自动驾驶中的关键环节,对道路安全与效率至关重要。然而,传统方法常受限于标注数据稀缺,在多智能体预测中表现不佳。为此,本文提出一种解耦的上下文感知预训练框架DECAMP,突破了现有方法将表征学习与预训练任务耦合的局限。该框架将行为模式学习与潜在特征重构分离,优先保证动态可解释性,从而提升下游预测的场景表征能力。同时,结合上下文感知表示学习与协作式时空预训练任务,实现结构与意图推理的联合优化,捕捉潜在动态意图。在Argoverse 2基准测试中,本方法展现出卓越性能,验证了其在多智能体轨迹预测中的有效性。据我们所知,这是首个面向自动驾驶的多智能体轨迹预测上下文自编码框架。代码与模型将公开共享。
原文摘要 · Abstract (English)
Trajectory prediction is a critical component of autonomous driving, essential for ensuring both safety and efficiency on the road. However, traditional approaches often struggle with the scarcity of labeled data and exhibit suboptimal performance in multi-agent prediction scenarios. To address these challenges, we introduce a disentangled context-aware pre-training framework for multi-agent motion prediction, named DECAMP. Unlike existing methods that entangle representation learning with pretext tasks, our framework decouples behavior pattern learning from latent feature reconstruction, prioritizing interpretable dynamics and thereby enhancing scene representation for downstream prediction. Additionally, our framework incorporates context-aware representation learning alongside collaborative spatial-motion pretext tasks, which enables joint optimization of structural and intentional reasoning while capturing the underlying dynamic intentions. Our experiments on the Argoverse 2 benchmark showcase the superior performance of our method, and the results attained underscore its effectiveness in multi-agent motion forecasting. To the best of our knowledge, this is the first context autoencoder framework for multi-agent motion forecasting in autonomous driving. The code and models will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。