用语音模型思路提升人体姿态预测,更抗真实噪声。
Scriboora: Rethinking Human Pose Forecasting
- 类比语音理解,用大模型适配姿态预测任务
- 在真实噪声下性能下降,但可经无监督微调恢复
- 提供统一训练评估流程,解决复现难题
人体姿态预测基于历史观测预测未来姿态,在动作识别、自动驾驶和人机交互等领域有重要应用。本文评估了多种姿态预测算法在绝对姿态预测任务中的表现,揭示了诸多复现问题,并提出统一的训练与评估流程。通过类比语音理解任务,证明近期语音模型可高效迁移至姿态预测,显著提升现有最先进性能。最后,使用从姿态估计模型获取的含噪关节点坐标评估模型鲁棒性,引入新数据集变体,发现估计姿态导致性能显著下降,但可通过无监督微调部分恢复。
原文摘要 · Abstract (English)
Human pose forecasting predicts future poses based on past observations, and has many significant applications in areas such as action recognition, autonomous driving or human-robot interaction. This paper evaluates a wide range of pose forecasting algorithms in the task of absolute pose forecasting, revealing many reproducibility issues, and provides a unified training and evaluation pipeline. After drawing a high-level analogy to the task of speech understanding, it is shown that recent speech models can be efficiently adapted to the task of pose forecasting, and improve current state-of-the-art performance. Finally, the robustness of the models is evaluated, using noisy joint coordinates obtained from a pose estimation model, to reflect a realistic type of noise, which is closer to real-world applications. For this a new dataset variation is introduced, and it is shown that estimated poses result in a substantial performance degradation, and how much of it can be recovered again by unsupervised finetuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。