用大模型先验提升单帧压缩成像的3D重建质量与稳定性
GS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model Priors

- 结合大视觉模型先验与3D高斯点云,从单张压缩图像重建高质量3D场景
- 在多个基准上达到最优重建质量,对视角变化具有强鲁棒性
- 适合需要高速三维成像的科研与工业场景,如动态物体捕捉
快照压缩成像(SCI)通过将时空信息压缩至单个二维测量中,实现高速视频采集及相对运动下的多视角场景捕获。现有方法在3D重建中面临信息丢失、视角多样性有限及联合优化3D表示与相机位姿的计算负担等挑战。本文提出GS²CI框架,利用3D高斯溅射(3DGS)与大规模视觉基础模型(VFMs)的强先验能力,从单一SCI测量中重建高质量3D场景。核心方法包括:基于测量的3D VFM初始化与面向SCI的高斯优化;粗粒度收敛后,借助辅助2D VFM在合成视点提供伪视图监督以精修局部外观。为缓解SCI监督模糊导致的3DGS优化不稳定性,提出专用的密度化策略OSGR——通过局部透明度统计筛选分裂候选,均值透明度调控防止损失补偿式透明度膨胀,并施加显式候选比例与高斯数量约束以控制表达增长。大量实验表明,本方法在多个基准上综合表现最强,兼具领先的重建质量、视角变化鲁棒性与竞争力的计算效率。
原文摘要 · Abstract (English)
Snapshot Compressive Imaging (SCI) offers an efficient solution for high-speed video acquisition and, under exposure-time camera--scene relative motion, multi-view scene capture by compressing temporal or spatial information into a single 2D measurement. While recent studies have explored SCI for 3D scene reconstruction, existing methods struggle with significant challenges due to information loss, limited viewpoint diversity, and the computational burden of jointly optimizing 3D representations and camera poses. In this work, we propose a novel framework that reconstructs high-quality 3D scenes from a single SCI measurement by leveraging 3D Gaussian Splatting (3DGS) and the powerful priors of large-scale vision foundation models (VFMs). Our primary reconstruction combines measurement-derived 3D VFM initialization with SCI-aware Gaussian optimization. After coarse-stage convergence, an auxiliary 2D VFM provides pseudo-view supervision at synthesized viewpoints for local appearance refinement. To further address the instability caused by ambiguous SCI supervision during 3DGS optimization, we introduce Opacity-Guided Splitting and Growth Regulation (OSGR), an SCI-specific densification strategy that augments split candidates using local opacity statistics, discourages loss-compensating opacity inflation through mean-opacity regulation, and bounds representation growth with explicit candidate-ratio and Gaussian-count constraints. Extensive experiments across multiple benchmarks demonstrate that our method achieves the strongest overall performance, combining leading reconstruction quality and robustness to viewpoint variation with competitive computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。