让多视角图像生成更清晰3D场景,解决视图越多越模糊的问题
AVSplat: Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning

- 用精选辅助视图预处理,聚焦关键信息提升对应稳定性
- 视图增多时效果不降反升,单点图像质量显著提升
- 适合高精度3D重建、无标定多视角输入场景
无位姿约束的前馈3D高斯点云渲染可从未校准的多视角图像中合成新视图。尽管更多视图本应提升性能,但现有方法在密集视图输入下常因全局聚合注意力分散、朴素体素融合将大量高斯点平均为过度平滑表示而退化。本文提出AVSplat框架,将额外视图转化为可靠的聚合与表征信号。在全局注意力前,每个视图与一组精选的、具有相关性和多样性的辅助视图进行一次轻量级交互,缓存特征提供聚焦的场景上下文以稳定对应关系。在表征层面,采用自适应温度感知的体素融合策略,在高密度区域通过占用率和点置信度引导增强归属精度。关键优势在于恢复了正向视图扩展效应:随着输入视图增多,性能保持稳定或提升,而非在密集视图下退化。消融实验表明,辅助视图预处理主要防止密集视图退化,而占用率引导的体素融合贡献了最主要的单点图像质量提升。
原文摘要 · Abstract (English)
Pose-free feed-forward 3D Gaussian Splatting enables novel view synthesis from uncalibrated multi-view images. Although more views should improve performance, existing methods often degrade with dense-view inputs because global aggregation spreads attention over many tokens, and naive voxel fusion averages many Gaussians into overly smooth representations. We present AVSplat, a framework that turns additional views into reliable signals for both aggregation and representation. Before global attention, each view performs a single lightweight interaction with a small set of Assist Views chosen for relevance and diversity, and the cached features provide a focused scene context that stabilizes correspondence. For representation, we use adaptive temperature-aware voxel fusion that sharpens attribution under high occupancy, guided by occupancy and point confidence. Crucially, AVSplat restores positive view scaling where performance remains stable or improves as more input views are added, instead of degrading in the dense-view regime. Ablations show that Assist View Preconditioning is primarily responsible for preventing dense-view degradation, while Occupancy-guided Voxel Fusion contributes most of the single-point image-quality gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。