简单高效3D交互分割方法,跨场景性能领先
Easy3D: A Simple Yet Effective Method for 3D Interactive Segmentation
- 基于体素稀疏编码器+轻量Transformer解码器,实现点击融合
- 在ScanNet等5个数据集上超越现有最佳方法,尤其在未见场景表现优异
- 适合需要快速精准3D分割的机器人、VR/AR应用开发者
随着图像重建、生成和机器人扫描获取的数字3D环境日益普及,对3D交互的需求持续增长,例如3D交互分割,可用于物体选择与操作。现有方法在效率、精度及跨域泛化能力方面仍存不足,尤其面对未知对象或几何分布时表现不稳定。本文提出Easy3D,一种简单高效的3D交互分割方法,通过体素稀疏编码器与轻量级Transformer解码器结合,实现隐式点击融合。该方法在包含ScanNet、ScanNet++、S3DIS、KITTI-360在内的多个基准数据集上均超越当前最优模型,并在高斯泼溅(Gaussian Splatting)生成的未见几何分布数据上也表现出显著优势。项目主页:https://simonelli-andrea.github.io/easy3d。
原文摘要 · Abstract (English)
The increasing availability of digital 3D environments, whether through image-based 3D reconstruction, generation, or scans obtained by robots, is driving innovation across various applications. These come with a significant demand for 3D interaction, such as 3D Interactive Segmentation, which is useful for tasks like object selection and manipulation. Additionally, there is a persistent need for solutions that are efficient, precise, and performing well across diverse settings, particularly in unseen environments and with unfamiliar objects. In this work, we introduce a 3D interactive segmentation method that consistently surpasses previous state-of-the-art techniques on both in-domain and out-of-domain datasets. Our simple approach integrates a voxel-based sparse encoder with a lightweight transformer-based decoder that implements implicit click fusion, achieving superior performance and maximizing efficiency. Our method demonstrates substantial improvements on benchmark datasets, including ScanNet, ScanNet++, S3DIS, and KITTI-360, and also on unseen geometric distributions such as the ones obtained by Gaussian Splatting. The project web-page is available at https://simonelli-andrea.github.io/easy3d.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。