用大重建模型实现快速可控3D生成,避免2D-3D对齐问题。
ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model
- 基于大重建模型的端到端框架,直接生成3D内容。
- 在G-OBJ、GSO、ABO三数据集上均表现优异,生成质量高。
- 适合需要高效可控3D生成的研究与工业应用。
尽管3D生成方法取得进展,可控性仍是难题。现有基于分数蒸馏采样的方法耗时长,且先生成2D再映射至3D的过程缺乏内在对齐。为此,我们提出ControLRM,一种基于大重建模型(LRM)的端到端前馈模型,实现快速可控3D生成。ControLRM包含2D条件生成器、条件编码变压器和三平面解码器变压器。不从零训练,而是采用联合训练框架:在条件训练分支锁定三平面解码器,复用预训练于数百万3D数据的深层编码层;在图像训练分支解锁解码器,建立2D与3D表示间的隐式对齐。为确保评估无偏,样本来自G-OBJ、GSO、ABO三个独立数据集,而非人工挑选。定量与定性对比实验表明,该方法在3D可控性与生成质量方面具备强泛化能力。
原文摘要 · Abstract (English)
Despite recent advancements in 3D generation methods, achieving controllability still remains a challenging issue. Current approaches utilizing score-distillation sampling are hindered by laborious procedures that consume a significant amount of time. Furthermore, the process of first generating 2D representations and then mapping them to 3D lacks internal alignment between the two forms of representation. To address these challenges, we introduce ControLRM, an end-to-end feed-forward model designed for rapid and controllable 3D generation using a large reconstruction model (LRM). ControLRM comprises a 2D condition generator, a condition encoding transformer, and a triplane decoder transformer. Instead of training our model from scratch, we advocate for a joint training framework. In the condition training branch, we lock the triplane decoder and reuses the deep and robust encoding layers pretrained with millions of 3D data in LRM. In the image training branch, we unlock the triplane decoder to establish an implicit alignment between the 2D and 3D representations. To ensure unbiased evaluation, we curate evaluation samples from three distinct datasets (G-OBJ, GSO, ABO) rather than relying on cherry-picking manual generation. The comprehensive experiments conducted on quantitative and qualitative comparisons of 3D controllability and generation quality demonstrate the strong generalization capacity of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。