arXiv:2606.16638cs.CV2026-06

构建工业场景3D重建基准数据集,揭示当前方法在真实工业环境中的局限性。

MVM-IOD: An Industrial Object-Centric Benchmark Dataset for the Evaluation of 3D Reconstruction Methods

  • 用机械臂移动相机环绕物体采集9类工业品图像,生成18个场景
  • 发现前馈式方法在新拍摄条件下点云与位姿精度显著下降
  • 建议对输入图像做简单预处理以提升工业场景适用性

工业场景中的3D物体重建与相机位姿估计极具挑战,因误差代价高且计算时间受限。典型工业物体的复杂性进一步增加了难度。现有数据集多未反映真实工业场景。为此,我们提出机器视觉计量工业物体数据集(MVM-IOD),通过安装在工业机器人末端的相机在物体周围半球形轨迹采集图像。MVM-IOD包含9个物体、2种背景,共18个场景,提供参考相机位姿和3D点云,支持基于图像的3D重建、位姿估计及新视角生成方法的评估。基于该数据集,我们系统评测了当前主流方法,包括结构光从运动(SfM)、多视图立体(MVS)、近期前馈方法(如视觉几何接地变换器π3)以及2D高斯泼溅。实验表明,此类采集方式产生的图像分布偏离前馈模型训练分布,导致点云与位姿性能下降;但通过简单预处理可使图像分布更接近训练数据,从而提升性能。因此,在特定工业应用中应谨慎使用前馈方法。

原文摘要 · Abstract (English)

3D object reconstruction, and camera pose estimation in industrial applications are challenging tasks, as errors are costly while the computation time is often limited. The complexity of typical industrial objects further complicates these tasks. Most of the existing datasets in this context do not depict realistic industrial scenarios. Therefore, we introduce the Machine Vision Metrology Industrial Object Dataset (MVM-IOD). Images of typical industrial objects are captured systematically, by moving a camera, mounted at the end effector of an industrial robot arm, on a hemisphere around the objects. MVM-IOD contains reference camera poses and reference 3D point clouds, the acquired RGB images of 9 objects and 2 background choices resulting in 18 scenes, which allows evaluation of all image based methods that compute a 3D reconstruction, camera poses, or novel views of a scene. Based on MVM-IOD, we extensively evaluate current SOTA 3D reconstruction and camera pose estimation methods, such as Structure from Motion, Multi-View Stereo, recent feed forward methods (Visual Geometry Grounded Transformer, π3), and 2D Gaussian Splatting and report our findings as a baseline for future research. The experiments show that capture setups like ours generate out-of distribution images for feed forward methods, leading to suboptimal point clouds and camera poses. However, these out-of-distribution images can be shifted closer to the training distribution by applying simple preprocessing steps. Consequently, in certain industrial applications, feed forward methods should be used with caution.

3D重建工业视觉基准数据集位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。