用少量低清图像重建高清3D场景,靠参考图和纹理自适应提升细节。
SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images
- 结合外部参考图与内部纹理线索,增强低分辨率输入的细节信息。
- 在RealEstate10K等数据集上优于现有方法,跨数据集泛化能力强。
- 适合自动驾驶、机器人等需要快速高精度3D重建的场景。
从稀疏、低分辨率(LR)图像进行前馈式3D重建是自动驾驶和具身AI等实际应用的关键能力。然而,现有方法难以恢复精细纹理细节,根源在于低分辨率输入缺乏高频信息。为此,我们提出SRSplat,一个仅需少数低分辨率视图即可重建高分辨率3D场景的前馈框架。核心思想是通过联合利用外部高质量参考图像与内部纹理线索来弥补纹理信息不足。我们首先为每个场景构建特定参考图库,该图库由多模态大语言模型(MLLMs)和扩散模型生成。为融合外部信息,引入参考引导特征增强(RGFE)模块,对齐并融合低分辨率输入图像与其参考孪生图像的特征。随后,使用多视图融合特征训练解码器以预测高斯原型。为进一步优化预测的高斯原型,提出纹理感知密度控制(TADC)模块,根据输入图像的纹理丰富度自适应调整高斯密度。大量实验表明,SRSplat在RealEstate10K、ACID和DTU等多个数据集上均优于现有方法,并展现出强大的跨数据集与跨分辨率泛化能力。
原文摘要 · Abstract (English)
Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency information in LR inputs. To address this, we propose \textbf{SRSplat}, a feed-forward framework that reconstructs high-resolution 3D scenes from only a few LR views. Our main insight is to compensate for the deficiency of texture information by jointly leveraging external high-quality reference images and internal texture cues. We first construct a scene-specific reference gallery, generated for each scene using Multimodal Large Language Models (MLLMs) and diffusion models. To integrate this external information, we introduce the \textit{Reference-Guided Feature Enhancement (RGFE)} module, which aligns and fuses features from the LR input images and their reference twin image. Subsequently, we train a decoder to predict the Gaussian primitives using the multi-view fused feature obtained from \textit{RGFE}. To further refine predicted Gaussian primitives, we introduce \textit{Texture-Aware Density Control (TADC)}, which adaptively adjusts Gaussian density based on the internal texture richness of the LR inputs. Extensive experiments demonstrate that our SRSplat outperforms existing methods on various datasets, including RealEstate10K, ACID, and DTU, and exhibits strong cross-dataset and cross-resolution generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。