arXiv:2411.17945cs.CVcs.AI2024-11CVPR被引 7

构建4000万条文本描述的3D数据集,提升文本生成高质量3D模型的能力。

MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

  • 用多阶段自动化流程生成从详细描述到标签的多层次文本注释。
  • 在GPT-4和人工评估中分别取得72.41%和73.40%的胜率,优于现有数据集。
  • 适合需要高保真3D生成与快速原型设计的研究者与开发者。

由于现有数据集规模小、多样性不足且标注深度有限,从文本提示生成高保真3D内容仍是计算机视觉中的重大挑战。为此,我们推出了MARVEL-40M+,一个包含超过890万3D资产和4000万条文本注释的大型数据集,整合自七个主要3D数据集。我们的贡献在于一种新颖的多阶段注释流程,结合开源预训练多视角视觉语言模型(VLM)和大语言模型(LLM),自动生成从详细描述(150–200词)到简洁语义标签(10–20词)的多层次文本。该结构支持细粒度3D重建与快速原型设计。此外,我们将源数据集的人类元数据引入注释流程,增强领域特定信息并减少VLM幻觉。同时,我们开发了MARVEL-FX3D,一种两阶段文本到3D生成流程:先用注释微调Stable Diffusion,再通过预训练图像到3D网络在15秒内生成带纹理的3D网格。大量评估表明,MARVEL-40M+在注释质量和语言多样性方面显著优于现有数据集,其在GPT-4评估中获得72.41%胜率,在人工评估中达到73.40%。

原文摘要 · Abstract (English)

Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.

3D生成文本生成数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。