用像素空间统一3D生成与重建,提升真实感和几何精度。
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

- 直接在渲染图像上做扩散,避免隐空间编码损失。
- 引入几何感知损失,使多视角对齐真实三维结构。
- 单模型兼顾生成与重建,性能超越主流方法。
3D重建与生成通常采用不同范式:重建基于像素回归,生成依赖隐空间扩散。现有统一方法在隐空间中定义扩散目标,导致优化偏离真实3D表示,且因隐编码引入信息损失,需预训练变分自编码器(VAE)或表征自编码器(RAE)。本文提出PixWorld,将两类任务统一于像素空间扩散框架下,通过直接在渲染图像上监督扩散过程,消除上述局限,并使优化更贴近3D场景保真度。除了传统的光度与感知监督外,还引入几何感知损失,利用预训练3D基础模型的几何感知特征空间,对齐渲染视图与真实视图的几何结构,提供3D结构监督。PixWorld在多个指标上优于现有隐空间生成方法,达到顶尖重建方法水平,验证了统一像素空间方法的优势。
原文摘要 · Abstract (English)
3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。