arXiv:2605.18262cs.RO2026-05

用条件变分自编码器提升行人轨迹预测的多样性与准确性

On Improving Multimodal Pedestrian Trajectory Prediction with CVAE: A Study on Benchmark and Robot Data

论文配图:On Improving Multimodal Pedestrian Trajectory Prediction with CVAE: A Study on Benchmark and Robot Data
图 1 · 摘自论文原文
  • 在Social-STGCNN基础上引入CVAE建模多模态未来轨迹
  • 在真实机器人数据上实现更优终点精度和轨迹多样性
  • 适合自动驾驶与配送机器人在复杂环境中的轨迹预测

准确的行人轨迹预测对自主系统在复杂环境(如模块化公交车、郊区或半结构化区域的配送机器人)中运行至关重要。社交时空图卷积神经网络(Social-STGCNN)通过建模社会交互展现了强劲性能,但生成多样化且校准良好的未来轨迹仍具挑战。本文在Social-STGCNN主干基础上,引入基于条件变分自编码器(CVAE)的概率建模方法,显式捕捉多模态未来轨迹。我们在ETH和UCY行人轨迹数据集以及移动机器人采集的真实世界行人数据集上评估该方法。结果表明,在公开基准上取得适度提升,但在不同人群配置下展现出更一致的终点精度和更好的轨迹多样性。机器人采集数据上的评估进一步证明该方法在非精心筛选基准下的有效性,支持其在实际部署中的应用。

原文摘要 · Abstract (English)

Accurate pedestrian trajectory prediction is crucial for autonomous systems operating in complex environments, such as modular buses and delivery robots in suburban or semi-structured areas. Social Spatio-Temporal Graph Convolutional Neural Networks (Social-STGCNN) have shown strong performance by modeling social interactions; however, producing diverse and well-calibrated future trajectories remains challenging. In this work, we build on a Social-STGCNN backbone and introduce a Conditional Variational Autoencoder (CVAE)-based probabilistic formulation to explicitly model multimodal future trajectories. We evaluate the method on the ETH and UCY pedestrian trajectory datasets as well as on a real-world pedestrian dataset collected by a mobile robot. Results show moderate gains on public benchmarks, but more consistent endpoint accuracy and improved trajectory diversity across different crowd configurations. Evaluation on robot-collected data further demonstrates the approach's effectiveness beyond curated benchmarks and supports its applicability in practical deployments.

轨迹预测多模态机器人CVAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。