提出可高效扩展的视图合成模型,显著降低训练成本。
Scaling View Synthesis Transformers
- 采用编码器-解码器架构,实现计算资源最优利用。
- 在多个算力水平下表现优于现有方法,且训练耗时大幅减少。
- 适合追求高效训练与高性能视图合成的研究者使用。
无几何建模的视图合成变压器近期在新视角合成(NVS)任务中达到领先性能,超越依赖显式几何建模的传统方法。然而其算力扩展规律尚不明确。本文系统研究了视图合成变压器的缩放定律,提出了训练算力最优的NVS模型设计原则。与先前结论相反,我们发现编码器-解码器架构可实现算力最优;早期负面结果源于次优架构选择及训练算力不均的对比。在多个算力水平下,我们提出的可扩展视图合成模型(SVSM)表现出与仅解码器模型相当的扩展性,实现了更优的性能-算力帕累托前沿,并在真实世界NVS基准上超越此前最先进水平,同时显著降低训练算力需求。
原文摘要 · Abstract (English)
Geometry-free view synthesis transformers have recently achieved state-of-the-art performance in Novel View Synthesis (NVS), outperforming traditional approaches that rely on explicit geometry modeling. Yet the factors governing their scaling with compute remain unclear. We present a systematic study of scaling laws for view synthesis transformers and derive design principles for training compute-optimal NVS models. Contrary to prior findings, we show that encoder-decoder architectures can be compute-optimal; we trace earlier negative results to suboptimal architectural choices and comparisons across unequal training compute budgets. Across several compute levels, we demonstrate that our encoder-decoder architecture, which we call the Scalable View Synthesis Model (SVSM), scales as effectively as decoder-only models, achieves a superior performance-compute Pareto frontier, and surpasses the previous state-of-the-art on real-world NVS benchmarks with substantially reduced training compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。