用少图生成不拘网格的3D高斯,重建更完整、细节更清晰。
Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction
- 通过高斯参数的流匹配生成非网格对齐的3D高斯
- 仅用4张图即实现高质量重建,未观察区域无空洞或模糊
- 适合需要少量视图、高保真重建的场景
我们提出Free-Range Gaussians,一种从最少四张图像生成非像素、非体素对齐3D高斯的多视角重建方法。该方法基于高斯参数的流匹配,以生成式建模方式实现非网格对齐数据的监督,并能在未观测区域合成合理内容。相比先前依赖冗余网格对齐高斯的方法,本方法避免了孔洞和条件均值模糊问题。为处理高质量重建所需的高数量高斯,我们引入分层分块机制,将空间相关的高斯聚合成联合变换器标记,序列长度减半且保持结构完整性。训练中采用时间步加权渲染损失,推理时结合光度梯度引导与无分类器引导以提升保真度。在Objaverse和Google Scanned Objects数据集上的实验表明,该方法在使用显著更少高斯的情况下,持续优于像素和体素对齐方法,尤其在输入视图缺失物体部分时表现更优。
原文摘要 · Abstract (English)
We present Free-Range Gaussians, a multi-view reconstruction method that predicts non-pixel, non-voxel-aligned 3D Gaussians from as few as four images. This is done through flow matching over Gaussian parameters. Our generative formulation of reconstruction allows the model to be supervised with non-grid-aligned 3D data, and enables it to synthesize plausible content in unobserved regions. Thus, it improves on prior methods that produce highly redundant grid-aligned Gaussians, and suffer from holes or blurry conditional means in unobserved regions. To handle the number of Gaussians needed for high-quality results, we introduce a hierarchical patching scheme to group spatially related Gaussians into joint transformer tokens, halving the sequence length while preserving structure. We further propose a timestep-weighted rendering loss during training, and photometric gradient guidance and classifier-free guidance at inference to improve fidelity. Experiments on Objaverse and Google Scanned Objects show consistent improvements over pixel and voxel-aligned methods while using significantly fewer Gaussians, with large gains when input views leave parts of the object unobserved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。