用单图生成高质量3D图像,解决视角一致性难题。
F3D-Gaus: Feed-forward 3D-aware Generation on ImageNet with Cycle-Aggregative Gaussian Splatting

- 基于像素对齐高斯点云的前馈生成框架,实现单图3D重建。
- 在ImageNet上生成多视角一致、细节丰富的3D图像,质量接近2D顶尖水平。
- 引入循环聚合约束与视频先验,提升3D纹理与视角泛化能力,适合3D内容生成研究者。
本文针对从单视图数据集(如ImageNet)中进行可泛化的3D感知生成问题。核心挑战是在无多视图或动态数据情况下学习鲁棒的3D感知表征,并保证不同视角间纹理与几何的一致性。尽管已有基线方法能实现3D感知生成,但其生成图像质量仍远逊于顶尖2D生成模型。为此,我们提出一种新颖的前馈式生成流水线F3D-Gaus,基于像素对齐高斯点云,可从单视图输入生成更真实可靠的3D渲染结果。同时,引入自监督循环聚合约束,强化所学3D表征的跨视角一致性。该训练策略天然支持多个对齐高斯原语的聚合,显著缓解了单视图像素对齐高斯点云固有的插值局限。此外,通过引入视频模型先验,实现几何感知优化,提升了宽视角场景下精细细节的生成能力,增强了模型捕捉复杂3D纹理的能力。大量实验表明,本方法不仅实现了从单视图数据集出发的高质量、多视角一致的3D感知生成,还显著提升了训练与推理效率。
原文摘要 · Abstract (English)
This paper tackles the problem of generalizable 3D-aware generation from monocular datasets, e.g., ImageNet. The key challenge of this task is learning a robust 3D-aware representation without multi-view or dynamic data, while ensuring consistent texture and geometry across different viewpoints. Although some baseline methods are capable of 3D-aware generation, the quality of the generated images still lags behind state-of-the-art 2D generation approaches, which excel in producing high-quality, detailed images. To address this severe limitation, we propose a novel feed-forward pipeline based on pixel-aligned Gaussian Splatting, coined as F3D-Gaus, which can produce more realistic and reliable 3D renderings from monocular inputs. In addition, we introduce a self-supervised cycle-aggregative constraint to enforce cross-view consistency in the learned 3D representation. This training strategy naturally allows aggregation of multiple aligned Gaussian primitives and significantly alleviates the interpolation limitations inherent in single-view pixel-aligned Gaussian Splatting. Furthermore, we incorporate video model priors to perform geometry-aware refinement, enhancing the generation of fine details in wide-viewpoint scenarios and improving the model's capability to capture intricate 3D textures. Extensive experiments demonstrate that our approach not only achieves high-quality, multi-view consistent 3D-aware generation from monocular datasets, but also significantly improves training and inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。