arXiv:2502.12546cs.CV2025-02中稿 · ICRA被引 5

用自由移动的人体动作实现多摄像头时空标定,无需标记物。

Spatiotemporal Multi-Camera Calibration using Freely Moving People

  • 将人体姿态转为单位球面点,统一求解旋转、时间偏移和匹配关系。
  • 通过软分配与共面约束,联合优化多视角时空对齐与对应关系。
  • 适用于真实场景,无需标记,适合无标记多摄像头系统部署。

我们提出一种利用多视角视频中自由移动人群进行时空多相机标定的新方法。由于多相机标定与视图间匹配问题本质相互依赖,将其统一建模面临重大挑战。本文将该问题视为两组3D点的配准问题,利用动态多人场景中的人体运动特性。具体地,采用现成的单目3D人体姿态估计算法获取3D人体姿态,并将其转换为单位球面上的3D点,交替求解旋转、时间偏移与关联关系。采用概率方法,通过视图间的软分配联合求解时空数据对齐与对应关系建立。平面对齐约束用于确定平移参数。将成对注册结果整合至多视角框架,并通过非线性优化提升相机位姿、时间偏移及多人关联的精度。在合成数据与真实数据上的大量实验表明,该方法有效且灵活,可作为实用的无标记标定工具。

原文摘要 · Abstract (English)

We propose a novel method for spatiotemporal multi-camera calibration using freely moving people in multiview videos. Since calibrating multiple cameras and finding matches across their views are inherently interdependent, performing both in a unified framework poses a significant challenge. We address these issues as a single registration problem of matching two sets of 3D points, leveraging human motion in dynamic multi-person scenes. To this end, we utilize 3D human poses obtained from an off-the-shelf monocular 3D human pose estimator and transform them into 3D points on a unit sphere, to solve the rotation, time offset, and the association alternatingly. We employ a probabilistic approach that can jointly solve both problems of aligning spatiotemporal data and establishing correspondences through soft assignment between two views. The translation is determined by applying coplanarity constraints. The pairwise registration results are integrated into a multiview setup, and then a nonlinear optimization method is used to improve the accuracy of the camera poses, temporal offsets, and multi-person associations. Extensive experiments on synthetic and real data demonstrate the effectiveness and flexibility of the proposed method as a practical marker-free calibration tool.

多相机标定人体姿态时空对齐无标记

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。