让3D生成与感知互相提升,一次训练同时搞定真实场景生成和语义预测。
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

- 用语义占据引导扩散模型,联合训练实现生成与感知双向增强。
- 在室内室外场景中生成逼真3D场景,同时提升语义占据预测准确率。
- 适合需要高质量3D生成与感知协同的自动驾驶、机器人研究者。
近期扩散模型在3D场景生成与感知任务中表现优异,但现有方法通常将两者分离,仅作为下游感知任务的数据增强手段。本文提出OccScene,一种统一框架下的跨任务互学习新范式,实现细粒度3D感知与高质量生成的深度融合,达成双向增益。OccScene仅依赖文本提示生成新颖且一致的3D真实场景,通过联合训练扩散框架中的语义占据进行引导。为对齐占据与扩散隐空间,引入基于Mamba的双路径对齐模块,融合细粒度语义与几何先验。在该框架下,感知模块可借助定制化、多样化的生成场景得到有效提升,而感知先验反过来也增强生成质量,形成良性循环。大量实验表明,OccScene可在广泛室内与室外场景中实现高保真3D场景生成,同时显著提升3D感知任务中语义占据预测性能。
原文摘要 · Abstract (English)
Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two processes, acting as a data augmenter to generate synthetic data for downstream perception tasks. In this work, we propose OccScene, a novel mutual learning paradigm that integrates fine-grained 3D perception and high-quality generation in a unified framework, achieving a cross-task win-win effect. OccScene generates new and consistent 3D realistic scenes only depending on text prompts, guided with semantic occupancy in a joint-training diffusion framework. To align the occupancy with the diffusion latent, a Mamba-based Dual Alignment module is introduced to incorporate fine-grained semantics and geometry as perception priors. Within OccScene, the perception module can be effectively improved with customized and diverse generated scenes, while the perception priors in return enhance the generation performance for mutual benefits. Extensive experiments show that OccScene achieves realistic 3D scene generation in broad indoor and outdoor scenarios, while concurrently boosting the perception models to achieve substantial performance improvements in the 3D perception task of semantic occupancy prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。