提出动态图建模方法,提升多智能体未来轨迹预测准确性
ProgD: Progressive Multi-scale Decoding with Dynamic Graphs for Joint Multi-agent Motion Forecasting
- 用动态异构图逐步建模交互演化过程
- 在INTERACTION和Argoverse 2上均排名第一
- 适合自动驾驶场景中多车交互预测任务
准确预测周边智能体的运动对自动驾驶车辆的安全规划至关重要。近年来,预测方法已从单个智能体扩展到多个相互作用智能体的联合预测,但现有方法忽视了交互关系随时间演变的特性。为此,本文提出一种名为ProgD的渐进式多尺度解码策略,结合动态异构图进行场景建模。通过逐步展开动态异构图,设计因子化架构以处理时空依赖性,并逐步消除多智能体未来运动的不确定性。同时引入多尺度解码机制,提升未来场景建模与运动预测的一致性。所提方法在INTERACTION多智能体预测基准和Argoverse 2多世界预测基准上均取得领先性能,排名首位。
原文摘要 · Abstract (English)
Accurate motion prediction of surrounding agents is crucial for the safe planning of autonomous vehicles. Recent advancements have extended prediction techniques from individual agents to joint predictions of multiple interacting agents, with various strategies to address complex interactions within future motions of agents. However, these methods overlook the evolving nature of these interactions. To address this limitation, we propose a novel progressive multi-scale decoding strategy, termed ProgD, with the help of dynamic heterogeneous graph-based scenario modeling. In particular, to explicitly and comprehensively capture the evolving social interactions in future scenarios, given their inherent uncertainty, we design a progressive modeling of scenarios with dynamic heterogeneous graphs. With the unfolding of such dynamic heterogeneous graphs, a factorized architecture is designed to process the spatio-temporal dependencies within future scenarios and progressively eliminate uncertainty in future motions of multiple agents. Furthermore, a multi-scale decoding procedure is incorporated to improve on the future scenario modeling and consistent prediction of agents' future motion. The proposed ProgD achieves state-of-the-art performance on the INTERACTION multi-agent prediction benchmark, ranking $1^{st}$, and the Argoverse 2 multi-world forecasting benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。