用扩散模型实现一键控场景生成,无需反复试错。
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
- 两阶段框架:先编码场景,再根据遮罩条件生成
- 在SemanticKITTI和CarlaSC上生成质量优于现有方法
- 可生成未见数据集的新城市场景,适合快速建模
基于3D扩散模型的语义场景生成近期受到关注。然而,现有方法依赖无条件生成,编辑场景时需多次重采样,严重限制可控性与灵活性。为此,我们提出SSEditor,一种可控的语义场景编辑器,可在不重复采样的情况下生成指定目标类别。SSEditor采用两阶段扩散框架:(1) 训练3D场景自编码器以获得潜在三平面特征;(2) 训练遮罩条件扩散模型实现可定制的3D语义场景生成。第二阶段引入几何-语义融合模块,增强模型对几何与语义信息的学习能力,确保物体生成位置、大小与类别正确。在SemanticKITTI和CarlaSC上的大量实验表明,SSEditor在目标生成的可控性与灵活性、语义场景生成与重建质量方面均优于先前方法。更重要的是,在未见的Occ-3D Waymo数据集上的实验显示,SSEditor具备生成新城市场景的能力,支持3D场景的快速构建。
原文摘要 · Abstract (English)
Recent advancements in 3D diffusion-based semantic scene generation have gained attention. However, existing methods rely on unconditional generation and require multiple resampling steps when editing scenes, which significantly limits their controllability and flexibility. To this end, we propose SSEditor, a controllable Semantic Scene Editor that can generate specified target categories without multiple-step resampling. SSEditor employs a two-stage diffusion-based framework: (1) a 3D scene autoencoder is trained to obtain latent triplane features, and (2) a mask-conditional diffusion model is trained for customizable 3D semantic scene generation. In the second stage, we introduce a geometric-semantic fusion module that enhance the model's ability to learn geometric and semantic information. This ensures that objects are generated with correct positions, sizes, and categories. Extensive experiments on SemanticKITTI and CarlaSC demonstrate that SSEditor outperforms previous approaches in terms of controllability and flexibility in target generation, as well as the quality of semantic scene generation and reconstruction. More importantly, experiments on the unseen Occ-3D Waymo dataset show that SSEditor is capable of generating novel urban scenes, enabling the rapid construction of 3D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。