arXiv:2506.20367cs.GRcs.CV2025-06被引 1

用文字生成全景3D场景,支持自由编辑和沉浸浏览。

DreamAnywhere: Object-Centric Panoramic 3D Scene Generation

  • 分步生成全景图,分离背景与物体,再构建3D场景。
  • 生成的场景在多视角下连贯性更强,图像质量达到顶尖水平。
  • 适合低成本影视制作快速迭代场景布局与视觉风格。

文本到3D场景生成技术近年来展现出巨大潜力,可变革多个行业的内容创作。尽管研究取得了显著进展,现有方法通常仅生成前向视角,视觉保真度低,场景理解有限,且常针对室内或室外环境单独优化。本文提出DreamAnywhere,一个模块化系统,用于快速生成和原型设计3D场景。该系统从文本合成360°全景图,分解为背景与物体,通过混合修复构建完整3D表示,并将物体掩码提升为详细3D对象,放置于虚拟环境中。系统支持沉浸式导航与直观的对象级编辑,适用于场景探索、视觉预览和快速原型设计,几乎无需手工建模。特别适合低预算电影制作,可快速迭代场景布局与视觉基调。模块化设计允许独立替换组件。相比当前最先进的文本与图像驱动3D场景生成方法,DreamAnywhere在新视角合成的一致性上显著提升,图像质量具有竞争力,适用于多样且复杂的场景。用户研究表明,参与者明显更偏好本方法,验证了其技术鲁棒性与实用性。

原文摘要 · Abstract (English)

Recent advances in text-to-3D scene generation have demonstrated significant potential to transform content creation across multiple industries. Although the research community has made impressive progress in addressing the challenges of this complex task, existing methods often generate environments that are only front-facing, lack visual fidelity, exhibit limited scene understanding, and are typically fine-tuned for either indoor or outdoor settings. In this work, we address these issues and propose DreamAnywhere, a modular system for the fast generation and prototyping of 3D scenes. Our system synthesizes a 360° panoramic image from text, decomposes it into background and objects, constructs a complete 3D representation through hybrid inpainting, and lifts object masks to detailed 3D objects that are placed in the virtual environment. DreamAnywhere supports immersive navigation and intuitive object-level editing, making it ideal for scene exploration, visual mock-ups, and rapid prototyping -- all with minimal manual modeling. These features make our system particularly suitable for low-budget movie production, enabling quick iteration on scene layout and visual tone without the overhead of traditional 3D workflows. Our modular pipeline is highly customizable as it allows components to be replaced independently. Compared to current state-of-the-art text and image-based 3D scene generation approaches, DreamAnywhere shows significant improvements in coherence in novel view synthesis and achieves competitive image quality, demonstrating its effectiveness across diverse and challenging scenarios. A comprehensive user study demonstrates a clear preference for our method over existing approaches, validating both its technical robustness and practical usefulness.

3D生成全景场景文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。