arXiv:2507.07519cs.CV2025-07

构建首个多视角动态视频物体分割数据集,推动3D分割研究发展

MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

  • 基于真实场景采集17个动态场景,每场景9-46个视角
  • 包含7830张带4D运动标注的图像,覆盖459个实例、73类物体
  • 提供新基准与评估指标,适配多视角视频分割研究者

基于神经辐射场(NeRF)和3D高斯喷溅(3D GS)的方法在静态场景3D物体分割中日益流行,适用于多种3D场景理解与编辑任务。然而,动态场景的4D物体分割因缺乏足够规模且标注精确的多视角视频数据集而研究不足。本文提出MUVOD,一个用于训练与评估重建真实场景中物体分割的新多视角视频数据集。所选17个场景涵盖不同室内外活动,数据来自多种相机阵列来源,每场景至少9视图,最多46视图。共提供7830张RGB图像(每视频30帧),及其对应的4D运动分割掩码,支持同一视角内或同相机阵列跨视角的目标跟踪。该数据集包含459个实例、73个类别,旨在作为多视角视频分割方法的基础评估基准。我们还提出一种评估指标与基线分割方法,以推动该领域进展。此外,从MUVOD中选取50个不同条件、不同场景的物体,构建新的3D物体分割基准,支持对前沿3D分割方法的全面分析。MUVOD数据集可于https://volumetric-repository.labs.b-com.com/#/muvod获取。

原文摘要 · Abstract (English)

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of 3D scene understanding and editing tasks. Nevertheless, the 4D object segmentation of dynamic scenes remains an underexplored field due to the absence of a sufficiently extensive and accurately labelled multi-view video dataset. In this paper, we present MUVOD, a new multi-view video dataset for training and evaluating object segmentation in reconstructed real-world scenarios. The 17 selected scenes, describing various indoor or outdoor activities, are collected from different sources of datasets originating from various types of camera rigs. Each scene contains a minimum of 9 views and a maximum of 46 views. We provide 7830 RGB images (30 frames per video) with their corresponding segmentation mask in 4D motion, meaning that any object of interest in the scene could be tracked across temporal frames of a given view or across different views belonging to the same camera rig. This dataset, which contains 459 instances of 73 categories, is intended as a basic benchmark for the evaluation of multi-view video segmentation methods. We also present an evaluation metric and a baseline segmentation approach to encourage and evaluate progress in this evolving field. Additionally, we propose a new benchmark for 3D object segmentation task with a subset of annotated multi-view images selected from our MUVOD dataset. This subset contains 50 objects of different conditions in different scenarios, providing a more comprehensive analysis of state-of-the-art 3D object segmentation methods. Our proposed MUVOD dataset is available at https://volumetric-repository.labs.b-com.com/#/muvod.

视频分割3D分割多视角数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。