arXiv:2511.15565cs.CV2025-11

用语音模型思路提升人体姿态预测,更抗真实噪声。

Scriboora: Rethinking Human Pose Forecasting

  • 类比语音理解,用大模型适配姿态预测任务
  • 在真实噪声下性能下降,但可经无监督微调恢复
  • 提供统一训练评估流程,解决复现难题

人体姿态预测基于历史观测预测未来姿态,在动作识别、自动驾驶和人机交互等领域有重要应用。本文评估了多种姿态预测算法在绝对姿态预测任务中的表现,揭示了诸多复现问题,并提出统一的训练与评估流程。通过类比语音理解任务,证明近期语音模型可高效迁移至姿态预测,显著提升现有最先进性能。最后,使用从姿态估计模型获取的含噪关节点坐标评估模型鲁棒性,引入新数据集变体,发现估计姿态导致性能显著下降,但可通过无监督微调部分恢复。

原文摘要 · Abstract (English)

Human pose forecasting predicts future poses based on past observations, and has many significant applications in areas such as action recognition, autonomous driving or human-robot interaction. This paper evaluates a wide range of pose forecasting algorithms in the task of absolute pose forecasting, revealing many reproducibility issues, and provides a unified training and evaluation pipeline. After drawing a high-level analogy to the task of speech understanding, it is shown that recent speech models can be efficiently adapted to the task of pose forecasting, and improve current state-of-the-art performance. Finally, the robustness of the models is evaluated, using noisy joint coordinates obtained from a pose estimation model, to reflect a realistic type of noise, which is closer to real-world applications. For this a new dataset variation is introduced, and it is shown that estimated poses result in a substantial performance degradation, and how much of it can be recovered again by unsupervised finetuning.

姿态预测语音模型鲁棒性迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。