arXiv:2603.21064cs.CV2026-03被引 1

将几何与外观分离建模,用两专家结构实现快速3D高斯点云生成。

2Xplat: Decoupling Geometry and Appearance Modeling for Feed-Forward 3D Gaussian Splatting

  • 采用双专家设计,先预测相机位姿再生成3D高斯,解耦几何与外观建模。
  • 仅用不到5000次训练迭代,性能媲美有位姿的先进方法。
  • 模块化设计突破统一架构局限,适合追求高效3D重建的研究者。

无位姿约束的前向3D高斯点云(3DGS)为快速3D建模开辟了新方向,可在单次前向传播中从未校准的多视角图像生成高质量高斯表示。现有主流方法通常采用统一的单体架构,基于以几何为中心的3D基础模型,在单一网络中联合估计相机位姿并合成3DGS表示,导致几何推理与外观建模在共享表征中纠缠。本文提出2Xplat,一种基于双专家设计的无位姿前向3DGS框架,显式分离几何估计与高斯生成:专用几何专家首先预测相机位姿,再由外观专家据此合成3D高斯。尽管概念简单且此前研究较少关注,该双专家流水线在少于5000次训练迭代下即超越先前无位姿前向3DGS方法,性能达到当前最先进有位姿方法水平。结果挑战了统一架构的主流范式,表明模块化设计在复杂3D几何估计与外观合成任务中具有潜在优势。

原文摘要 · Abstract (English)

Pose-free feed-forward 3D Gaussian Splatting (3DGS) has opened a new frontier for rapid 3D modeling, enabling high-quality Gaussian representations to be generated from uncalibrated multi-view images in a single forward pass. The dominant approach adopts unified monolithic architectures, often built on geometry-centric 3D foundation models, to jointly estimate camera poses and synthesize 3DGS representations within a single network, entangling geometric reasoning and appearance modeling within a shared representation. In this work, we introduce 2Xplat, a pose-free feed-forward 3DGS framework based on a two-experts design that explicitly separates geometry estimation from Gaussian generation: a dedicated geometry expert first predicts camera poses, which are then passed to an appearance expert that synthesizes 3D Gaussians. Despite its conceptual simplicity, and being largely underexplored in prior works, our two-experts pipeline outperforms prior pose-free feed-forward 3DGS approaches in fewer than 5K training iterations, achieving performance on par with state-of-the-art posed methods. These results challenge the prevailing unified paradigm and suggest the potential advantages of modular design for complex 3D geometric estimation and appearance synthesis tasks.

3D高斯点云生成双专家前向建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。