用合成数据提升3D实例分割的泛化能力
ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
- 从海量3D模型中随机选物,增强几何与场景多样性
- 结合大模型推理与搜索算法,生成合理物体布局
- 多视角渲染融合生成真实感点云,适合训练新模型
无类别3D实例分割需在不依赖语义类别的情况下分割所有物体实例,包括未见类别。现有方法因标注3D场景数据稀少或2D分割噪声而泛化能力差。尽管合成数据生成有潜力,但现有3D场景合成方法难以同时满足几何多样性、上下文复杂性和布局合理性三大需求。为此,我们提出ASSIST-3D——一种专为无类别3D实例分割设计的自适应场景合成流程。其核心创新包括:1)从大规模3D CAD资产库中异构选物,通过采样随机性最大化几何与上下文多样性;2)基于大语言模型引导的空间推理与深度优先搜索,生成合理物体布局;3)通过多视角RGB-D图像渲染与融合,构建接近真实传感器数据的点云。在ScanNetV2、ScanNet++和S3DIS基准上的实验表明,使用ASSIST-3D生成数据训练的模型显著优于现有方法,验证了本方案在提升模型泛化能力方面的有效性。
原文摘要 · Abstract (English)
Class-agnostic 3D instance segmentation tackles the challenging task of segmenting all object instances, including previously unseen ones, without semantic class reliance. Current methods struggle with generalization due to the scarce annotated 3D scene data or noisy 2D segmentations. While synthetic data generation offers a promising solution, existing 3D scene synthesis methods fail to simultaneously satisfy geometry diversity, context complexity, and layout reasonability, each essential for this task. To address these needs, we propose an Adapted 3D Scene Synthesis pipeline for class-agnostic 3D Instance SegmenTation, termed as ASSIST-3D, to synthesize proper data for model generalization enhancement. Specifically, ASSIST-3D features three key innovations, including 1) Heterogeneous Object Selection from extensive 3D CAD asset collections, incorporating randomness in object sampling to maximize geometric and contextual diversity; 2) Scene Layout Generation through LLM-guided spatial reasoning combined with depth-first search for reasonable object placements; and 3) Realistic Point Cloud Construction via multi-view RGB-D image rendering and fusion from the synthetic scenes, closely mimicking real-world sensor data acquisition. Experiments on ScanNetV2, ScanNet++, and S3DIS benchmarks demonstrate that models trained with ASSIST-3D-generated data significantly outperform existing methods. Further comparisons underscore the superiority of our purpose-built pipeline over existing 3D scene synthesis approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。