arXiv:2602.00463cs.CV2026-02中稿 · ICASSP2026被引 4

用文本生成高保真全景3D场景,解决多视角不一致问题

PSGS: Text-driven Panorama Sliding Scene Generation via Gaussian Splatting

  • 分两阶段优化:先解析文本语义布局,再通过大模型反馈细化视觉细节
  • 采用全景滑动机制初始化3D点云,结合深度与语义损失提升一致性
  • 适合需要高效生成沉浸式内容的VR/AR/游戏开发者使用

从文本生成真实感3D场景对虚拟现实、增强现实和游戏等沉浸式应用至关重要。尽管文本驱动方法具有效率优势,但现有方法受限于3D-文本数据稀缺及多视角拼接不一致,导致场景过于简单。为此,我们提出PSGS,一种两阶段高保真全景场景生成框架。首先,创新的双层优化架构生成语义一致的全景图:布局推理层将文本解析为结构化空间关系,自优化层通过迭代多模态大模型反馈精修视觉细节。其次,全景滑动机制通过战略性采样重叠视角,初始化全局一致的3D高斯点云。训练中引入深度与语义一致性损失,显著提升渲染场景的质量与细节保真度。实验表明,PSGS在全景生成上优于现有方法,并生成更具吸引力的3D场景,为可扩展沉浸式内容创作提供可靠解决方案。

原文摘要 · Abstract (English)

Generating realistic 3D scenes from text is crucial for immersive applications like VR, AR, and gaming. While text-driven approaches promise efficiency, existing methods suffer from limited 3D-text data and inconsistent multi-view stitching, resulting in overly simplistic scenes. To address this, we propose PSGS, a two-stage framework for high-fidelity panoramic scene generation. First, a novel two-layer optimization architecture generates semantically coherent panoramas: a layout reasoning layer parses text into structured spatial relationships, while a self-optimization layer refines visual details via iterative MLLM feedback. Second, our panorama sliding mechanism initializes globally consistent 3D Gaussian Splatting point clouds by strategically sampling overlapping perspectives. By incorporating depth and semantic coherence losses during training, we greatly improve the quality and detail fidelity of rendered scenes. Our experiments demonstrate that PSGS outperforms existing methods in panorama generation and produces more appealing 3D scenes, offering a robust solution for scalable immersive content creation.

3D生成文本生成全景高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。