用更精确的骨骼表示减少姿态预测翻转,提升真实场景下的准确性。
Generalized Pose Space Embeddings for Training In-the-Wild using Anaylis-by-Synthesis
- 采用语义化的骨骼表征区分左右,避免传统方法的翻转问题。
- 在标准基准上显著降低翻转率,提升三维姿态预测精度。
- 适合需要高精度姿态估计的真实世界应用,如视频分析与动作识别。
当前姿态估计模型依赖大规模人工标注数据集,成本高昂且难以覆盖真实世界中的人体姿态与外观多样性。随着神经渲染和基于分析-合成框架的发展,可通过生成图像来训练模型,从而减少对人工标注的依赖。然而,现有方法因使用简化的中间骨骼表示,导致预测结果存在大量翻转,影响精度并阻碍后续三维定位等任务。为此,本文提出更具表现力的中间骨骼表示,能够捕捉姿态的语义信息(如左右区分),显著减少翻转现象。为有效训练该表示,我们扩展了分析-合成框架,引入基于合成数据的训练协议。实验表明,新方法在标准基准上优于以往基于分析-合成训练的模型,大幅降低翻转率并提升预测准确性。
原文摘要 · Abstract (English)
Modern pose estimation models are trained on large, manually-labelled datasets which are costly and may not cover the full extent of human poses and appearances in the real world. With advances in neural rendering, analysis-by-synthesis and the ability to not only predict, but also render the pose, is becoming an appealing framework, which could alleviate the need for large scale manual labelling efforts. While recent work have shown the feasibility of this approach, the predictions admit many flips due to a simplistic intermediate skeleton representation, resulting in low precision and inhibiting the acquisition of any downstream knowledge such as three-dimensional positioning. We solve this problem with a more expressive intermediate skeleton representation capable of capturing the semantics of the pose (left and right), which significantly reduces flips. To successfully train this new representation, we extend the analysis-by-synthesis framework with a training protocol based on synthetic data. We show that our representation results in less flips and more accurate predictions. Our approach outperforms previous models trained with analysis-by-synthesis on standard benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。