用3D一致性约束提升稀疏视角三维重建质量
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

- 基于高质量对应点构建3D属性一致性损失
- 在ScanNet++上比Splatt3R平均提升4.5dB以上PSNR
- 适合需要高保真重建的稀疏视角场景应用
从稀疏视角重建高保真3D场景是通用神经渲染的核心挑战。现有通用3D高斯泼溅(3DGS)方法在稀疏视角下常出现几何伪影,因仅依赖2D光度损失无法解决深度与对应关系模糊问题。为此,我们提出MAC-Splat,一种基于直接3D一致性监督的训练框架。该方法依托MASt3R几何主干网络和冻结的DINOv3编码器获取语义感知的2D对应关系,作为3D监督的几何锚点。利用这些锚点,定义多属性一致性(MAC)损失,通过统一世界坐标系强制匹配高斯点的位置、形状和外观的一致性,该公式对异常值鲁棒且尊重协方差矩阵几何结构,确保稀疏视角下稳定训练。在ScanNet++上的实验表明,MAC-Splat优于强基线模型,尤其在不同重叠率下表现显著,平均PSNR相比Splatt3R提升超过4.5 dB,LPIPS降低,且随相机位姿差距增大仍保持性能。结果表明,结合高质量对应关系的直接多属性3D一致性目标,有效应对稀疏视角重建的病态问题。
原文摘要 · Abstract (English)
Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。