1秒生成360°全景3D高精度重建,速度比传统方法快800倍。
Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian Splats
- 混合Mamba2与Transformer结构,配合轻量合并与剪枝提升效率
- 输入960x540分辨率32张图,单卡A1001秒完成重建
- 支持大场景、多视角,兼容2D高斯等变体,适合工业级3D重建
我们提出Long-LRM,一种前馈式3D高斯重建模型,可实现即时、高分辨率、360°全覆盖的场景级重建。该模型接收32张960×540分辨率图像,在单张A100 GPU上仅需1秒即可生成高斯表示。为应对大规模输入(25万词元)带来的长序列挑战,Long-LRM结合最新的Mamba2块与经典Transformer块,并引入轻量级标记合并模块和高斯剪枝步骤,在质量和效率间取得平衡。在大型DL3DV基准和Tanks&Temples数据集上评估显示,其重建质量媲美基于优化的方法,同时相较后者提速800倍,且输入规模较前人前馈方法至少扩大60倍。我们进行了广泛的消融实验,验证模型设计对渲染质量与计算效率的影响。此外,还探索了Long-LRM与2D GS等高斯变体的兼容性,显著提升了几何重建能力。
原文摘要 · Abstract (English)
We propose Long-LRM, a feed-forward 3D Gaussian reconstruction model for instant, high-resolution, 360° wide-coverage, scene-level reconstruction. Specifically, it takes in 32 input images at a resolution of 960x540 and produces the Gaussian reconstruction in just 1 second on a single A100 GPU. To handle the long sequence of 250K tokens brought by the large input size, Long-LRM features a mixture of the recent Mamba2 blocks and the classical transformer blocks, enhanced by a light-weight token merging module and Gaussian pruning steps that balance between quality and efficiency. We evaluate Long-LRM on the large-scale DL3DV benchmark and Tanks&Temples, demonstrating reconstruction quality comparable to the optimization-based methods while achieving an 800x speedup w.r.t. the optimization-based approaches and an input size at least 60x larger than the previous feed-forward approaches. We conduct extensive ablation studies on our model design choices for both rendering quality and computation efficiency. We also explore Long-LRM's compatibility with other Gaussian variants such as 2D GS, which enhances Long-LRM's ability in geometry reconstruction. Project page: https://arthurhero.github.io/projects/llrm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。