用加速采样提升轨迹预测的多模态能力和效率。
Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling
- 采用DDIM采样,推理速度提升100倍
- 生成多条轨迹并用高斯混合模型表示不确定性
- 适合需要高置信度多路径预测的自动驾驶场景
准确且具备不确定性感知的轨迹预测仍是自动驾驶的核心挑战,源于复杂的多智能体交互、多样的场景上下文以及未来运动的固有随机性。基于扩散的生成模型最近在捕捉多模态未来方面展现出强大潜力,但现有方法如cVMD存在采样慢、生成多样性利用不足和场景编码脆弱等问题。本文提出cVMDx,一种增强的基于扩散的轨迹预测框架,提升了效率、鲁棒性和多模态预测能力。通过DDIM采样,cVMDx实现推理时间最高降低100倍,支持实用的多样本生成以估计不确定性。进一步使用拟合的高斯混合模型从生成轨迹中获得可解析的多模态预测。同时评估了基于CVQ-VAE的场景编码方案。在公开的highD数据集上的实验表明,cVMDx相比cVMD具有更高的准确性与显著的效率提升,实现了完全随机的多模态轨迹预测。
原文摘要 · Abstract (English)
Accurate and uncertainty-aware trajectory prediction remains a core challenge for autonomous driving, driven by complex multi-agent interactions, diverse scene contexts and the inherently stochastic nature of future motion. Diffusion-based generative models have recently shown strong potential for capturing multimodal futures, yet existing approaches such as cVMD suffer from slow sampling, limited exploitation of generative diversity and brittle scenario encodings. This work introduces cVMDx, an enhanced diffusion-based trajectory prediction framework that improves efficiency, robustness and multimodal predictive capability. Through DDIM sampling, cVMDx achieves up to a 100x reduction in inference time, enabling practical multi-sample generation for uncertainty estimation. A fitted Gaussian Mixture Model further provides tractable multimodal predictions from the generated trajectories. In addition, a CVQ-VAE variant is evaluated for scenario encoding. Experiments on the publicly available highD dataset show that cVMDx achieves higher accuracy and significantly improved efficiency over cVMD, enabling fully stochastic, multimodal trajectory prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。