arXiv:2508.16577cs.CVcs.AI2025-08

用2D图片检索增强3D生成,让罕见概念更准确

MV-RAG: Retrieval Augmented Multiview Diffusion

  • 先从海量2D图库中检索相关图像,再用多视角扩散模型生成3D
  • 在罕见概念上3D一致性提升显著,生成更真实且贴合文本
  • 适合需要精准生成稀有物体的场景,如艺术设计、工业建模

文本到3D生成方法通过利用预训练的2D扩散先验取得了显著进展,能够生成高质量且3D一致的结果。然而,这些方法在处理域外(OOD)或罕见概念时往往表现不佳,导致结果不一致或不准确。为此,我们提出MV-RAG,一种新型文本到3D生成流程:首先从大规模真实世界2D数据库中检索相关2D图像,然后将多视角扩散模型基于这些图像进行条件生成,以合成一致且准确的多视角输出。该检索增强模型通过一种新颖的混合训练策略实现,融合结构化多视角数据与多样化2D图像集合。具体包括:在多视角数据上训练时,使用模拟检索差异的增强视图进行视图特异性重建;在检索到的真实世界2D图像集上,采用独特的保留视图预测目标——模型从其他视图预测被遮挡的视图,从而从2D数据中推断3D一致性。为支持严格的域外评估,我们引入一个新的具有挑战性的域外提示数据集。与当前最先进的文本到3D、图像到3D及个性化基线相比,我们的方法在罕见概念上显著提升了3D一致性、照片真实感和文本贴合度,同时在标准基准上保持了竞争力。

原文摘要 · Abstract (English)

Text-to-3D generation approaches have advanced significantly by leveraging pretrained 2D diffusion priors, producing high-quality and 3D-consistent outputs. However, they often fail to produce out-of-domain (OOD) or rare concepts, yielding inconsistent or inaccurate results. To this end, we propose MV-RAG, a novel text-to-3D pipeline that first retrieves relevant 2D images from a large in-the-wild 2D database and then conditions a multiview diffusion model on these images to synthesize consistent and accurate multiview outputs. Training such a retrieval-conditioned model is achieved via a novel hybrid strategy bridging structured multiview data and diverse 2D image collections. This involves training on multiview data using augmented conditioning views that simulate retrieval variance for view-specific reconstruction, alongside training on sets of retrieved real-world 2D images using a distinctive held-out view prediction objective: the model predicts the held-out view from the other views to infer 3D consistency from 2D data. To facilitate a rigorous OOD evaluation, we introduce a new collection of challenging OOD prompts. Experiments against state-of-the-art text-to-3D, image-to-3D, and personalization baselines show that our approach significantly improves 3D consistency, photorealism, and text adherence for OOD/rare concepts, while maintaining competitive performance on standard benchmarks.

3D生成多视角检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。