用2D图像作中介,实现文本驱动的高质量3D场景生成
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
- 通过生成2D图像作为中间步骤,引导3D模型构建
- 在布局与美学上优于现有方法,用户评测胜率74.89%
- 无需训练,适配多种风格,适合创意设计与快速原型
传统3D场景设计需艺术素养与复杂软件操作。尽管文本到3D生成技术已简化流程,但受限于高质量3D数据稀缺,性能常受制约。相比之下,基于网络图像训练的文本到图像模型能生成多样且可靠的2D空间布局与视觉风格。本文提出ArtiScene,一种无需训练的自动化流程:先根据场景描述生成2D图像,再从中提取物体形状与外观生成3D模型,并利用同一图像中的几何、位置与姿态信息组装最终场景。该方法可广泛适用于多种场景与风格,在定量指标上显著超越当前最优基准。在大规模用户测试中平均胜率达74.89%,在GPT-4o评估中达95.07%。
原文摘要 · Abstract (English)
Designing 3D scenes is traditionally a challenging task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have greatly simplified this process by letting users create scenes based on simple text descriptions. However, as these methods generally require extra training or in-context learning, their performance is often hindered by the limited availability of high-quality 3D data. In contrast, modern text-to-image models learned from web-scale images can generate scenes with diverse, reliable spatial layouts and consistent, visually appealing styles. Our key insight is that instead of learning directly from 3D scenes, we can leverage generated 2D images as an intermediary to guide 3D synthesis. In light of this, we introduce ArtiScene, a training-free automated pipeline for scene design that integrates the flexibility of free-form text-to-image generation with the diversity and reliability of 2D intermediary layouts. First, we generate 2D images from a scene description, then extract the shape and appearance of objects to create 3D models. These models are assembled into the final scene using geometry, position, and pose information derived from the same intermediary image. Being generalizable to a wide range of scenes and styles, ArtiScene outperforms state-of-the-art benchmarks by a large margin in layout and aesthetic quality by quantitative metrics. It also averages a 74.89% winning rate in extensive user studies and 95.07% in GPT-4o evaluation. Project page: https://artiscene-cvpr.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。