用旋转数据增强提升单目3D人体姿态估计的旋转鲁棒性
On the Role of Rotation Equivariance in Monocular 2D-to-3D Human Pose Lifting
- 通过输入输出同步旋转增强,学习旋转等变性
- 旋转后误差降低超30%,最高达72%
- 比设计上完全等变的方法快37倍,适合实际部署
我们研究单目3D人体姿态估计中2D到3D姿态提升模型在旋转下的性能退化问题。系统对比了不同具备旋转等变归纳偏置的架构,探究是否应通过网络结构强制等变性,或由数据学习。基于常见HPE基准测试,发现对输入和输出姿态联合应用旋转数据增强,可有效学习旋转等变性。相比无增强模型,旋转后误差降低超过30%,最高达72%;且该方法误差比完全设计等变的方法低最多15%,推理速度提升达37倍。结果在真实世界全身体旋转序列及先进模型上均具泛化能力。
原文摘要 · Abstract (English)
We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2D keypoint detection followed by 2D-to-3D lifting. Despite their success, we find that current lifting models exhibit strong performance degradation under rotations. We systematically study rotation equivariance across lifting architectures with increasing degrees of equivariant inductive bias, and investigate whether equivariance should be enforced by architectural design or learned from data, given the inherent ambiguity of monocular 2D-to-3D lifting. Utilising common HPE benchmarks, we demonstrate that rotation equivariance can be effectively learned via rotation-based data augmentation applied jointly to input and output poses. Compared with non-augmented models, this reduces error on rotated poses by over 30% across the benchmarks, with reductions of up to 72%. Moreover, augmentation-based models achieve up to 15% lower error than methods that are fully equivariant by design, while providing up to 37$\times$ faster inference. We further show that these findings generalise to real-world human pose sequences involving full-body rotations and to a state-of-the-art HPE model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。