通过逐步生成模式提升交通行为预测的多样性与准确性。
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
- 将多模态轨迹建模为序列,分步预测未来行为模式。
- 在多个基准上实现更高轨迹多样性与良好精度平衡。
- 无需后处理即可自然扩展模式数量,适合高不确定性场景。
预判未来事件的多模态性是保障自动驾驶安全的基础。然而,交通参与者的行为多模态预测因缺乏多模态真实标签而受限。现有方法多采用胜者通吃训练策略,仍存在轨迹多样性不足和模式置信度不准的问题。部分方法虽通过生成大量候选轨迹缓解此问题,但需依赖启发式后处理筛选最具代表性的模式,该过程缺乏统一原则且影响轨迹精度。为此,我们提出 ModeSeq,一种将模式视为序列进行建模的新范式。不同于一次性解码多个合理轨迹,ModeSeq 要求运动解码器逐步推断下一个模式,从而更显式地捕捉模式间的相关性,显著增强对多模态的推理能力。基于序列模式预测的归纳偏置,我们进一步提出早期匹配胜者通吃(EMTA)训练策略,进一步提升轨迹多样性。无需密集模式预测或启发式后处理,ModeSeq 显著提升多模态输出多样性,同时保持优异轨迹精度,在多个运动预测基准上表现均衡。此外,ModeSeq 天然具备模式外推能力,可在未来高度不确定时预测更多行为模式。
原文摘要 · Abstract (English)
Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this challenge, yet still suffer from limited trajectory diversity and uncalibrated mode confidence. While some approaches address these limitations by generating excessive trajectory candidates, they necessitate a post-processing stage to identify the most representative modes, a process lacking universal principles and compromising trajectory accuracy. We are thus motivated to introduce ModeSeq, a new multimodal prediction paradigm that models modes as sequences. Unlike the common practice of decoding multiple plausible trajectories in one shot, ModeSeq requires motion decoders to infer the next mode step by step, thereby more explicitly capturing the correlation between modes and significantly enhancing the ability to reason about multimodality. Leveraging the inductive bias of sequential mode prediction, we also propose the Early-Match-Take-All (EMTA) training strategy to diversify the trajectories further. Without relying on dense mode prediction or heuristic post-processing, ModeSeq considerably improves the diversity of multimodal output while attaining satisfactory trajectory accuracy, resulting in balanced performance on motion prediction benchmarks. Moreover, ModeSeq naturally emerges with the capability of mode extrapolation, which supports forecasting more behavior modes when the future is highly uncertain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。