arXiv:2501.10462cs.CVcs.AI2025-01AAAI被引 5

用文本或图像生成高质量3D场景,体积小且结构清晰。

BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

论文配图:BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
图 1 · 摘自论文原文
  • 分阶段构建点云并用3D高斯喷溅生成场景
  • 多层级深度先验正则化提升几何连续性
  • 结构化上下文压缩降低存储开销,适合虚拟现实

随着虚拟现实应用的普及,3D场景生成成为新兴研究前沿。3D场景结构复杂,需保证输出密集、连贯并包含完整结构。现有方法依赖预训练文本到图像扩散模型和单目深度估计器,但生成场景占用存储大,缺乏有效正则化,易产生几何畸变。为此,我们提出BloomScene,一种轻量级结构化3D高斯喷溅框架,可从文本或图像输入生成多样且高质量的3D场景。具体地,设计跨模态渐进式场景生成框架,通过增量点云重建与3D高斯喷溅实现连贯生成;提出基于分层深度先验的正则化机制,利用多级深度精度与平滑性约束增强真实感与连续性;最终引入结构化上下文引导压缩机制,借助结构化哈希网格建模无序锚点属性上下文,显著消除结构冗余,降低存储开销。多场景综合实验表明,该框架在多个基线中展现显著潜力与优势。

原文摘要 · Abstract (English)

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods rely on pre-trained text-to-image diffusion models and monocular depth estimators. However, the generated scenes occupy large amounts of storage space and often lack effective regularisation methods, leading to geometric distortions. To this end, we propose BloomScene, a lightweight structured 3D Gaussian splatting for crossmodal scene generation, which creates diverse and high-quality 3D scenes from text or image inputs. Specifically, a crossmodal progressive scene generation framework is proposed to generate coherent scenes utilizing incremental point cloud reconstruction and 3D Gaussian splatting. Additionally, we propose a hierarchical depth prior-based regularization mechanism that utilizes multi-level constraints on depth accuracy and smoothness to enhance the realism and continuity of the generated scenes. Ultimately, we propose a structured context-guided compression mechanism that exploits structured hash grids to model the context of unorganized anchor attributes, which significantly eliminates structural redundancy and reduces storage overhead. Comprehensive experiments across multiple scenes demonstrate the significant potential and advantages of our framework compared with several baselines.

3D生成高斯喷溅跨模态轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。