arXiv:2511.17609cs.CV2025-11

用多视角2D标注生成精准3D真实值,支持完整物体形状重建。

3D Ground Truth Reconstruction from Multi-Camera Annotations Using UKF

  • 基于无迹卡尔曼滤波融合多相机2D标注,实现3D坐标估计。
  • 在CMC、Wildtrack、Panoptic数据集上实现高精度3D定位,优于已有方法。
  • 无需3D标注,自动处理遮挡,适合多摄像机系统部署。

精确的3D真实值估计对自动驾驶、监控和机器人应用至关重要。本文提出一种新方法,利用无迹卡尔曼滤波(UKF)融合多台校准相机的2D边界框或姿态关键点标注,生成准确的3D真实值。通过利用人工标注的2D真值,所提出的多相机单目标跟踪算法,结合基于单应性的投影与UKF融合,将2D图像坐标转化为鲁棒的3D世界坐标。该算法处理多视角数据,估计物体位置与形状,有效应对遮挡等挑战。我们在CMC、Wildtrack和Panoptic数据集上进行评估,结果显示其3D定位精度显著优于现有3D真实值。与仅提供地面平面信息的方法不同,本方法还输出每个物体的完整3D形状。此外,该算法为仅使用2D图像标注的多摄像机系统提供了可扩展且全自动的解决方案。

原文摘要 · Abstract (English)

Accurate 3D ground truth estimation is critical for applications such as autonomous navigation, surveillance, and robotics. This paper introduces a novel method that uses an Unscented Kalman Filter (UKF) to fuse 2D bounding box or pose keypoint ground truth annotations from multiple calibrated cameras into accurate 3D ground truth. By leveraging human-annotated ground-truth 2D, our proposed method, a multi-camera single-object tracking algorithm, transforms 2D image coordinates into robust 3D world coordinates through homography-based projection and UKF-based fusion. Our proposed algorithm processes multi-view data to estimate object positions and shapes while effectively handling challenges such as occlusion. We evaluate our method on the CMC, Wildtrack, and Panoptic datasets, demonstrating high accuracy in 3D localization compared to the available 3D ground truth. Unlike existing approaches that provide only ground-plane information, our method also outputs the full 3D shape of each object. Additionally, the algorithm offers a scalable and fully automatic solution for multi-camera systems using only 2D image annotations.

3D重建多相机卡尔曼滤波目标跟踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。