用3D视觉大模型提升无姿态3D渲染质量
AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
- 训练时引入自洽位姿对齐,解决视角与几何不一致问题
- 通过评分机制筛选低质点云,提升重建清晰度
- 适合追求高质量无姿态3D渲染的研究者
尽管3D视觉基础模型(3DVFMs)在零样本视觉几何估计中表现优异,但其直接应用于通用新视角合成(NVS)仍具挑战。本文提出AirSplat,一种新训练框架,将3DVFMs的鲁棒几何先验有效融入高保真、无姿态的NVS。方法包含两项关键技术:(1) 自洽位姿对齐(SCPA),训练时反馈环确保像素级监督,消除位姿-几何偏差;(2) 基于评分的不透明度匹配(ROM),利用稀疏视图NVS教师模型的局部3D几何一致性知识,过滤劣化点元。大规模基准测试表明,该方法在重建质量上显著优于当前最先进的无姿态NVS方案。AirSplat展示了将3DVFMs用于同时实现视觉几何估计与高质量视角合成的巨大潜力。
原文摘要 · Abstract (English)
While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero-shot capabilities in visual geometry estimation, their direct application to generalizable novel view synthesis (NVS) remains challenging. In this paper, we propose AirSplat, a novel training framework that effectively adapts the robust geometric priors of 3DVFMs into high-fidelity, pose-free NVS. Our approach introduces two key technical contributions: (1) Self-Consistent Pose Alignment (SCPA), a training-time feedback loop that ensures pixel-aligned supervision to resolve pose-geometry discrepancy; and (2) Rating-based Opacity Matching (ROM), which leverages the local 3D geometry consistency knowledge from a sparse-view NVS teacher model to filter out degraded primitives. Experimental results on large-scale benchmarks demonstrate that our method significantly outperforms state-of-the-art pose-free NVS approaches in reconstruction quality. Our AirSplat highlights the potential of adapting 3DVFMs to enable simultaneous visual geometry estimation and high-quality view synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。