首个融合多模态的3D生成模型,用图像和点云协同提升生成质量。
Collaborative Multi-Modal Coding for High-Quality 3D Generation
- 通过协同编码融合图像、深度图和点云特征,保留各模态优势。
- 在小规模数据上实现媲美大模型的3D生成质量。
- 适合需要高质量3D内容生成的研究者与开发者使用。
3D内容天然具有多模态特性,可投影为不同模态(如RGB图像、RGBD和点云)。每种模态在3D资产建模中各有优势:RGB图像包含生动的3D纹理,而点云则定义精细的3D几何结构。然而,现有大多数3D原生生成架构主要依赖单一模态,忽略多模态互补优势;或仅限于3D结构,限制了可用训练数据集范围。为全面利用多模态信息进行3D建模,我们提出TriMM——首个前馈式3D原生生成模型,能够从基础多模态数据(如RGB、RGBD和点云)中学习。具体而言:1)引入协同多模态编码,整合模态特异性特征的同时保持其独特表示能力;2)引入辅助2D与3D监督,增强多模态编码的鲁棒性与性能;3)基于嵌入的多模态代码,采用三平面潜在扩散模型生成高质量3D资产,显著提升纹理与几何细节。在多个知名数据集上的实验表明,尽管训练数据量较小,TriMM仍能实现与大规模数据训练模型相当的竞争力表现。此外,我们在近期的RGB-D数据集上进行了额外实验,验证了将其他多模态数据集融入3D生成的可行性。
原文摘要 · Abstract (English)
3D content inherently encompasses multi-modal characteristics and can be projected into different modalities (e.g., RGB images, RGBD, and point clouds). Each modality exhibits distinct advantages in 3D asset modeling: RGB images contain vivid 3D textures, whereas point clouds define fine-grained 3D geometries. However, most existing 3D-native generative architectures either operate predominantly within single-modality paradigms-thus overlooking the complementary benefits of multi-modality data-or restrict themselves to 3D structures, thereby limiting the scope of available training datasets. To holistically harness multi-modalities for 3D modeling, we present TriMM, the first feed-forward 3D-native generative model that learns from basic multi-modalities (e.g., RGB, RGBD, and point cloud). Specifically, 1) TriMM first introduces collaborative multi-modal coding, which integrates modality-specific features while preserving their unique representational strengths. 2) Furthermore, auxiliary 2D and 3D supervision are introduced to raise the robustness and performance of multi-modal coding. 3) Based on the embedded multi-modal code, TriMM employs a triplane latent diffusion model to generate 3D assets of superior quality, enhancing both the texture and the geometric detail. Extensive experiments on multiple well-known datasets demonstrate that TriMM, by effectively leveraging multi-modality, achieves competitive performance with models trained on large-scale datasets, despite utilizing a small amount of training data. Furthermore, we conduct additional experiments on recent RGB-D datasets, verifying the feasibility of incorporating other multi-modal datasets into 3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。