arXiv:2412.01430cs.CVcs.AI2024-12被引 13

扩充至520万张多视角图像,构建更接近2D数据集规模的3D视觉基准。

MVImgNet2.0: A Larger-scale Dataset of Multi-view Images

论文配图:MVImgNet2.0: A Larger-scale Dataset of Multi-view Images
图 1 · 摘自论文原文
  • 将多视角图像扩展至520万张、515个类别,覆盖更广物体范围。
  • 新增360度拍摄支持完整重建,点云质量更高,相机位姿误差更低。
  • 适合3D生成、重建与视觉理解研究者使用,推动3D视觉发展。

MVImgNet2.0 是对原始 MVImgNet 的升级,涵盖约 520,000 个真实物体和 515 个类别,显著扩大了数据规模与类别覆盖范围,使其在规模上更接近 ImageNet 等主流 2D 数据集。新版本引入四项关键改进:(i) 多数物体采用 360° 多视角拍摄,支持更完整的 3D 重建;(ii) 采用更先进的分割方法,生成更高精度的前景掩码;(iii) 使用更强的结构光恢复(SfM)算法,降低相机位姿估计误差;(iv) 基于 360° 视角图像生成高质量稠密点云。大量实验验证了其在提升大型 3D 重建模型性能方面的有效性。数据集将公开发布于 luyues.github.io/mvimgnet2,包含所有 520,000 个物体的多视角图像、高保真点云及标注代码,旨在推动 3D 视觉研究发展。

原文摘要 · Abstract (English)

MVImgNet is a large-scale dataset that contains multi-view images of ~220k real-world objects in 238 classes. As a counterpart of ImageNet, it introduces 3D visual signals via multi-view shooting, making a soft bridge between 2D and 3D vision. This paper constructs the MVImgNet2.0 dataset that expands MVImgNet into a total of ~520k objects and 515 categories, which derives a 3D dataset with a larger scale that is more comparable to ones in the 2D domain. In addition to the expanded dataset scale and category range, MVImgNet2.0 is of a higher quality than MVImgNet owing to four new features: (i) most shoots capture 360-degree views of the objects, which can support the learning of object reconstruction with completeness; (ii) the segmentation manner is advanced to produce foreground object masks of higher accuracy; (iii) a more powerful structure-from-motion method is adopted to derive the camera pose for each frame of a lower estimation error; (iv) higher-quality dense point clouds are reconstructed via advanced methods for objects captured in 360-degree views, which can serve for downstream applications. Extensive experiments confirm the value of the proposed MVImgNet2.0 in boosting the performance of large 3D reconstruction models. MVImgNet2.0 will be public at luyues.github.io/mvimgnet2, including multi-view images of all 520k objects, the reconstructed high-quality point clouds, and data annotation codes, hoping to inspire the broader vision community.

3D视觉多视角数据集点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。