arXiv:2604.24312cs.CVcs.AI2026-04

无需相机标定即可精准估计多视角人体姿态,突破真实场景应用瓶颈。

Unconstrained Multi-view Human Pose Estimation with Algebraic Priors

论文配图:Unconstrained Multi-view Human Pose Estimation with Algebraic Priors
图 1 · 摘自论文原文
  • 用数据驱动的Transformer融合替代传统依赖标定参数的三角化方法。
  • 引入格罗布纳基损失,强制模型遵守投影几何规律,提升姿态准确性。
  • 利用人体运动的时序对称性,解决未标定场景下的尺度模糊问题。

从多视角图像中恢复3D人体姿态通常依赖精确的相机标定,但在真实场景中常不可得,严重限制了现有方法的应用。为此,我们提出一种无约束框架,结合深度神经网络、代数先验与时间动态,实现未标定多视角人体姿态估计。首先,提出三角化变换器(TTR),将经典三角化重构为数据驱动的令牌融合过程,摆脱对显式相机参数的依赖。其次,设计格罗布纳基校正器(GC),通过源自多视图代数簇的约束损失,确保神经预测严格遵循投影几何法则。最后,提出时序等变修正器(TER),利用人体运动的等变特性,施加时序一致性和结构一致性,有效缓解未标定设置下的尺度模糊问题。在标准基准上的大量实验表明,该框架在未标定多视角人体姿态估计上达到新SOTA,显著缩小了无标定方法与完全标定最优方案之间的性能差距。

原文摘要 · Abstract (English)

Recovering 3D human pose from multi-view imagery typically relies on precise camera calibration, which is often unavailable in real-world scenarios, thereby severely limiting the applicability of existing methods. To overcome this challenge, we propose an unconstrained framework that synergizes deep neural networks, algebraic priors, and temporal dynamics for uncalibrated multi-view human pose estimation. First, we introduce the Triangulation with Transformer Regressor (TTR), which reformulates classical triangulation into a data-driven token fusion process to bypass the dependency on explicit camera parameters. Second, to explicitly embed the inherent algebraic relations of the multi-view variety into the learning process, we propose the Gröbner basis Corrector (GC). This pioneering loss formulation enforces constraints derived from the multi-view variety to ensure the neural predictions strictly adhere to the laws of projective geometry. Finally, we devise the Temporal Equivariant Rectifier (TER), which exploits the equivariance property of human motion to impose temporal coherence and structural consistency, effectively mitigating scale ambiguity in uncalibrated settings. Extensive evaluations on standard benchmarks demonstrate that our framework establishes a new state-of-the-art for uncalibrated multi-view human pose estimation. Notably, our approach significantly closes the performance gap between calibration-free methods and fully calibrated oracles.

3D姿态估计多视角无标定代数先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。