arXiv:2609.07328cs.ROcs.AI2026-09

统一行人车辆异构运动预测,提升多智能体协同未来轨迹精度

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

论文配图:PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout
图 1 · 摘自论文原文
  • 基于历史轨迹的递归世界模型,同步建模行人根运动与关节动作、车辆状态
  • 在Waymo数据集上行人根位置误差降5.2%,关节误差降7.6%,车辆距离误差降11.9%
  • 单模型设计参数少57%,计算量降96.5%,适合实时部署的多智能体场景

行人-车辆局部预测涉及异构物理尺度:行人兼具整体位移与关节动作,车辆为刚体,由运动状态和朝向范围描述。现有方法通常忽略行人关节运动,而姿态预测器又不包含车辆未来。我们提出PV-WM,一种仅依赖历史感知轨迹的结构化世界模型。它递归推进行人根运动、15关节姿态及学习到的车辆状态,在同步异构状态中生成未来。生成的行人与车辆片段构成下一循环边界;车辆框通过预测中心与朝向重构,每次转移后重新计算行人-车辆几何关系。相较于匹配的一次性完整状态预测器,递归执行使根位置平均位移误差(Root ADE)降低12.7%,关键点误差(MPJPE)降低14.8%。在824个对齐的Waymo场景中(797个提供有效车辆未来),相比验证选中的模块化专家模型,PV-WM使根位置误差降低5.2%,关键点误差降7.6%,行人-车辆距离误差降11.9%,朝向框最近接近误差降5.8%。该单网络模型参数减少57.1%,每场景平均浮点运算量降低96.5%,测量延迟(p95)降低25.5%。PV-WM统一异构未来状态,同时保留行人与车辆特异性动力学。

原文摘要 · Abstract (English)

Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion with articulated motion, whereas vehicles are rigid bodies described by kinematic state and oriented extent. Existing road-agent forecasters typically omit pedestrian articulation, while pose forecasters leave vehicle futures outside the learned rollout. We introduce PV-WM, a history-only world model over structured post-perception tracks. It recurrently advances pedestrian root motion, 15-joint articulation, and learned vehicle states within a synchronized heterogeneous state. The generated pedestrian and vehicle chunks supply the next recurrent boundary; vehicle boxes are reconstructed from predicted center and heading with observed extent, and P-V geometry is recomputed after every transition. Relative to a matched one-shot complete-state predictor, recurrent execution reduces Root ADE by 12.7% and MPJPE by 14.8%. Feedback interventions show that later predictions depend on the content, temporal order, and pedestrian identity of generated articulation. Across 824 aligned Waymo contexts, with 797 providing valid future vehicle support, PV-WM reduces Root ADE by 5.2%, MPJPE by 7.6%, P-V distance error by 11.9%, and oriented-box closest-approach error by 5.8% relative to a validation-selected Modular Specialist. The single-network model uses 57.1% fewer parameters, 96.5% lower average FLOPs per local scene, and 25.5% lower measured p95 latency. PV-WM unifies this heterogeneous future state while preserving type-specific pedestrian and vehicle dynamics.

轨迹预测多智能体世界模型行人车辆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。