arXiv:2503.09950cs.CVcs.AI2025-03CVPR被引 55

用隐式最大似然蒸馏,实现100倍加速的多人轨迹预测。

MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation

  • 基于流匹配设计新损失函数,保证多模态轨迹准确且多样。
  • 在NBA、ETH-UCY等数据集上达到最先进性能。
  • 蒸馏后模型采样速度提升100倍,适合实时应用。

本文针对人类轨迹预测问题,旨在基于历史轨迹与上下文信息预测多人未来的多模态运动路径。提出新型条件流匹配模型MoFlow,可为场景中所有智能体生成K-shot未来轨迹。设计新颖的流匹配损失函数,确保至少一组预测轨迹准确,并鼓励所有K组轨迹保持多样性与合理性。结合隐式最大似然估计(IMLE),提出仅需教师模型样本即可完成蒸馏的新方法。在真实世界数据集SportVU NBA、ETH-UCY和SDD上的大量实验表明,教师模型与IMLE蒸馏的学生模型均达当前最优性能,能生成物理与社会合理的多样化轨迹。此外,学生模型采样速度比教师模型快100倍。

原文摘要 · Abstract (English)

In this paper, we address the problem of human trajectory forecasting, which aims to predict the inherently multi-modal future movements of humans based on their past trajectories and other contextual cues. We propose a novel motion prediction conditional flow matching model, termed MoFlow, to predict K-shot future trajectories for all agents in a given scene. We design a novel flow matching loss function that not only ensures at least one of the $K$ sets of future trajectories is accurate but also encourages all $K$ sets of future trajectories to be diverse and plausible. Furthermore, by leveraging the implicit maximum likelihood estimation (IMLE), we propose a novel distillation method for flow models that only requires samples from the teacher model. Extensive experiments on the real-world datasets, including SportVU NBA games, ETH-UCY, and SDD, demonstrate that both our teacher flow model and the IMLE-distilled student model achieve state-of-the-art performance. These models can generate diverse trajectories that are physically and socially plausible. Moreover, our one-step student model is $\textbf{100}$ times faster than the teacher flow model during sampling. The code, model, and data are available at our project page: https://moflow-imle.github.io

轨迹预测流匹配蒸馏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。