arXiv:2608.27039cs.CVcs.AI2026-08中稿 · GCPR 2026

建模人-物交互与社交关系,提升复杂场景下多人运动预测精度

Multi-Person Human Motion Forecasting in Complex Scenes

论文配图:Multi-Person Human Motion Forecasting in Complex Scenes
图 1 · 摘自论文原文
  • 用条件扩散模型融合运动历史、人际互动和物体线索
  • 在HiK和HOI-M3上2秒路径误差降低超30%
  • 支持多群体、多未来轨迹生成,适合智能驾驶与机器人场景

在复杂场景中准确预测多人运动,需对环境的过去与当前状态进行推理。在此背景下,如何有效整合物体信息与社交互动仍具挑战。为此,我们提出对象条件社交扩散(OCSD),一种将运动历史、多人交互和物体线索统一建模的条件扩散模型。OCSD采用对象条件机制,在每一步去噪时调节信号,实现细粒度的人-物交互推理;同时引入社交编码器,建模场景中所有人的相互作用。该模型自然适应不同规模的群体、复杂社交行为,并支持生成多个合理未来轨迹。大量实验表明,OCSD在Humans in Kitchens(HiK)和HOI-M3基准上均达到领先性能:相较于以往方法,其在HiK上2秒路径误差降低121.5 mm(31.3%),在HOI-M3上降低130.5 mm(33.2%),且长期预测更真实。

原文摘要 · Abstract (English)

Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.

运动预测扩散模型人-物交互多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。