arXiv:2606.29942cs.CV2026-06

基于视觉与姿态预测人类多种潜在行动目标,提升自主系统预判能力。

Scene-aware Prediction of Diverse Human Movement Goals

论文配图:Scene-aware Prediction of Diverse Human Movement Goals
图 1 · 摘自论文原文
  • 用条件变分自编码器融合场景图像和人体姿态,生成多样未来目标
  • 在GTA-IM和PROX数据集上实现跨场景泛化,可采样生成多条合理轨迹
  • 适用于需要预判人类行为的自动驾驶、机器人交互等场景

人类行为具有随机性,常由不同目标驱动。目标通常引导其运动,有助于长期轨迹预测。除个体社交线索外,环境上下文对推断人类意图至关重要。现有方法或需场景语义知识,或仅关注物体交互。本文提出一种基于生成模型的多目标预测新方法,利用当前RGB场景图像和人体姿态,通过条件变分自编码器(CVAE)预测多样化的未来运动目标。实验表明,该方法可通过采样CVAE隐空间生成多个合理的人类运动目标,在GTA-IM和PROX数据集上展现出良好泛化能力。代码已公开于https://github.com/Q-Y-Yang/DiverseGoalsPrediction.git。

原文摘要 · Abstract (English)

Anticipation of human behaviours facilitates autonomous systems in proactive planning. Human behaviour could be stochastic due to varying goals. Human goals typically guide their own movement and could therefore help to predict the human trajectory and human motion in the long-term. To infer the human movement intentions, the environmental context plays a significant role, in addition to the social cues expressed by the individual. Previous works on human goals prediction either require semantic knowledge of the scene, or only tackle interactions with objects. In this paper, we propose a novel multi-goal prediction method using the generative model to address the stochasticity of human movement. It leverages the current RGB scene and the human pose to predict diverse potential future goals of human movement based on the Conditional Variational Autoencoder (CVAE). Our results demonstrate that our approach is capable of generating multiple movement goals in the scene via samplings in latent space of the CVAE and exhibits generalization capability across scenarios in GTA-IM dataset and PROX dataset. Code is publicly available at \href{https://github.com/Q-Y-Yang/DiverseGoalsPrediction.git}{\texttt{https://github.com/Q-Y-Yang/DiverseGoalsPrediction}}.

行为预测生成模型多目标人体姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。