用扩散模型解决多视角人体姿态估计的标签效率难题
DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose Estimation

- 将多人身份分配建模为多项随机张量的生成扩散过程
- 仅用10%伪标签仍保持99%性能,显著提升标签效率
- 对相机布局不敏感,适合手术室等遮挡复杂场景
从多视角图像中恢复多人3D姿态是分析交互行为的基础瓶颈。现有自监督方法依赖合成姿态库,但因分布偏移导致真实场景泛化能力差。本文提出DisPOSE,将本质离散的多视角人物分配问题建模为多项随机张量上的生成扩散过程。通过在去噪过程中使用可微Sinkhorn投影,模型基于2D图像先验学习引导解向有效可行的分配。随后利用超图卷积解码器回归局部个体的完整3D骨架,显式建模跨视角的结构关系与关节连接。该方法在标准数据集上超越现有自监督最先进水平,并在新提出的手术室高度遮挡场景基准上表现优异。其扩散式定位在标签效率上表现出色,仅用10%伪标签即保持99%性能。解耦分配与根部回归模块且保持可微性,使DisPOSE近乎对不同相机布局无感。
原文摘要 · Abstract (English)
Recovering 3D human poses for multiple individuals from different camera views is a fundamental bottleneck for analyzing interacting behaviors. Existing self-supervised approaches leverage synthetic catalogues of 3D poses; however, this leads to poor generalization in real-world scenarios due to distribution shifts. We therefore introduce DisPOSE, a self-supervised framework that approximates the inherently discrete multi-view person-assignment problem as a generative diffusion process over the space of polystochastic tensors. By employing differentiable Sinkhorn projections during denoising, our model learns to guide solutions toward valid and feasible assignments based on 2D image priors. The complete 3D skeletons of localized individuals are then regressed using a Hypergraph-Convolutional Decoder that explicitly models relational structures and articulated joints across multiple views. The proposed approach outperforms current state-of-the-art self-supervised methods on standard datasets and demonstrates strong performance on a newly proposed benchmark featuring highly occluded scenes from surgical operating rooms. Our diffusion-based localization demonstrates high label efficiency, retaining 99% of its performance with only 10% of the pseudo-labels. Notably, disentangling the assignment and root regression components while maintaining differentiability makes DisPOSE nearly agnostic to different camera arrangements. Project site and code: https://wngtn.github.io/DisPOSE/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。