从一张图生成跨尺度3D世界,支持任意缩放细节
WonderZoom: Multi-Scale 3D World Generation
- 用自适应高斯面元实现多尺度3D内容生成与实时渲染
- 通过渐进式细节合成器逐级生成微观到宏观的结构
- 适合需要跨尺度3D重建的视觉设计与虚拟现实应用
我们提出WonderZoom,一种从单张图像生成跨空间尺度3D场景的新方法。现有3D世界生成模型仅限于单尺度合成,无法在不同粒度下保持场景一致性。根本挑战在于缺乏能生成和渲染显著差异空间尺寸内容的尺度感知3D表示。WonderZoom通过两项关键创新解决:(1) 尺度自适应高斯面元,用于生成与实时渲染多尺度3D场景;(2) 逐步细节合成器,迭代生成更精细尺度的3D内容。该方法使用户能够“缩放进入”3D区域,并自动回归生成此前不存在的细粒度细节,涵盖从景观到微观特征。实验表明,WonderZoom在质量与对齐度上显著优于现有视频与3D生成模型,实现了从单张图像生成多尺度3D世界。视频结果及交互式3D世界查看器见https://wonderzoom.github.io/
原文摘要 · Abstract (English)
We present WonderZoom, a novel approach to generating 3D scenes with contents across multiple spatial scales from a single image. Existing 3D world generation models remain limited to single-scale synthesis and cannot produce coherent scene contents at varying granularities. The fundamental challenge is the lack of a scale-aware 3D representation capable of generating and rendering content with largely different spatial sizes. WonderZoom addresses this through two key innovations: (1) scale-adaptive Gaussian surfels for generating and real-time rendering of multi-scale 3D scenes, and (2) a progressive detail synthesizer that iteratively generates finer-scale 3D contents. Our approach enables users to "zoom into" a 3D region and auto-regressively synthesize previously non-existent fine details from landscapes to microscopic features. Experiments demonstrate that WonderZoom significantly outperforms state-of-the-art video and 3D models in both quality and alignment, enabling multi-scale 3D world creation from a single image. We show video results and an interactive viewer of generated multi-scale 3D worlds in https://wonderzoom.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。