arXiv:2602.20160cs.CV2026-02中稿 · CVPR被引 8

用测试时训练实现长序列3D重建,线性计算量下高效生成高质量3D模型

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

  • 引入测试时训练层,将多视角图像压缩为隐空间快权重
  • 在物体与场景上均超越现有方法,收敛更快且质量更高
  • 适合需要持续更新的3D重建任务,如实时扫描与动态场景建模

我们提出tttLRM,一种新型大模型3D重建方法,通过测试时训练(TTT)层实现长上下文、自回归3D重建,计算复杂度为线性,进一步提升模型能力。该框架将多视角图像观测高效压缩至TTT层的快权重中,在隐空间形成隐式3D表示,并可解码为高斯点云(Gaussian Splats, GS)等显式格式,用于下游应用。模型的在线学习变体支持从流式观测中渐进式重建与优化。实验证明,预训练于新视角合成任务能有效迁移至显式3D建模,显著提升重建质量并加快收敛速度。大量实验表明,该方法在前馈式3D高斯重建上优于当前最先进方法,在物体与场景数据集上均取得更好性能。

原文摘要 · Abstract (English)

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability. Our framework efficiently compresses multiple image observations into the fast weights of the TTT layer, forming an implicit 3D representation in the latent space that can be decoded into various explicit formats, such as Gaussian Splats (GS) for downstream applications. The online learning variant of our model supports progressive 3D reconstruction and refinement from streaming observations. We demonstrate that pretraining on novel view synthesis tasks effectively transfers to explicit 3D modeling, resulting in improved reconstruction quality and faster convergence. Extensive experiments show that our method achieves superior performance in feedforward 3D Gaussian reconstruction compared to state-of-the-art approaches on both objects and scenes.

3D重建测试时训练自回归高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。