arXiv:2511.09502cs.CVcs.AI2025-11被引 2

用想象力生成3D人体姿态,让模型像人一样预测动作意图。

DreamPose3D: Hallucinative Diffusion with Prompt Learning for 3D Human Pose Estimation

  • 通过提示学习提取动作意图,动态引导去噪过程
  • 在Human3.6M和MPI-3DHP上达到最新最好性能
  • 适合处理模糊输入和复杂动作的鲁棒姿态估计

准确的3D人体姿态估计仍是关键挑战,需兼顾帧间时序一致性与关节关系的精细建模。现有方法多依赖几何线索独立预测每帧姿态,难以解决动作歧义并泛化至真实场景。受人类理解与预测运动方式启发,我们提出DreamPose3D,一种基于扩散模型的框架,融合动作感知推理与时序想象能力。该方法通过从2D姿态序列中提取任务相关的动作提示,动态调节去噪过程,捕捉高层意图。为有效建模关节间结构关系,引入融合运动学亲和性的表示编码器,嵌入注意力机制。最后,幻觉姿态解码器在训练中生成时序一致的3D姿态序列,模拟人类心理重建运动轨迹以化解感知歧义。在Human3.6M和MPI-3DHP等基准数据集上的实验表明,该方法在所有指标上均达到当前最优。进一步在广播棒球数据集上测试,即使面对模糊噪声的2D输入,仍表现出强鲁棒性,有效处理时序一致性和意图驱动的动作变化。

原文摘要 · Abstract (English)

Accurate 3D human pose estimation remains a critical yet unresolved challenge, requiring both temporal coherence across frames and fine-grained modeling of joint relationships. However, most existing methods rely solely on geometric cues and predict each 3D pose independently, which limits their ability to resolve ambiguous motions and generalize to real-world scenarios. Inspired by how humans understand and anticipate motion, we introduce DreamPose3D, a diffusion-based framework that combines action-aware reasoning with temporal imagination for 3D pose estimation. DreamPose3D dynamically conditions the denoising process using task-relevant action prompts extracted from 2D pose sequences, capturing high-level intent. To model the structural relationships between joints effectively, we introduce a representation encoder that incorporates kinematic joint affinity into the attention mechanism. Finally, a hallucinative pose decoder predicts temporally coherent 3D pose sequences during training, simulating how humans mentally reconstruct motion trajectories to resolve ambiguity in perception. Extensive experiments on benchmarked Human3.6M and MPI-3DHP datasets demonstrate state-of-the-art performance across all metrics. To further validate DreamPose3D's robustness, we tested it on a broadcast baseball dataset, where it demonstrated strong performance despite ambiguous and noisy 2D inputs, effectively handling temporal consistency and intent-driven motion variations.

3D姿态估计扩散模型动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。