arXiv:2505.21050cs.CV2025-05被引 4

用2.5D中间表示联合生成高质量3D形状与纹理。

Advancing high-fidelity 3D and Texture Generation with 2.5D latents

  • 构建统一的2.5D latent表示,融合多视角图像信息
  • 在文本和图像条件下生成高保真2.5D内容,支持结构与颜色一致
  • 轻量级2.5D转3D解码器,显著提升几何条件下的纹理生成质量

尽管大规模3D数据集和3D生成模型已有进展,但3D几何与纹理数据的复杂性和质量不均仍制约生成效果。现有方法通常分阶段使用不同模型生成几何与纹理,导致二者缺乏一致性。为此,我们提出一种联合生成3D几何与纹理的新框架。核心是生成可无缝转换的2.5D表示:首先将多视角RGB、法线与坐标图像整合为统一的2.5D latents;接着利用预训练2D基础模型,在文本和图像条件下游进行高保真2.5D生成;最后引入轻量级2.5D-to-3D refiner-decoder,高效生成细节丰富的3D表示。大量实验表明,本方法不仅能从文本和图像输入生成结构与色彩协调的高质量3D物体,且在几何条件下的纹理生成上显著优于现有方法。

原文摘要 · Abstract (English)

Despite the availability of large-scale 3D datasets and advancements in 3D generative models, the complexity and uneven quality of 3D geometry and texture data continue to hinder the performance of 3D generation techniques. In most existing approaches, 3D geometry and texture are generated in separate stages using different models and non-unified representations, frequently leading to unsatisfactory coherence between geometry and texture. To address these challenges, we propose a novel framework for joint generation of 3D geometry and texture. Specifically, we focus in generate a versatile 2.5D representations that can be seamlessly transformed between 2D and 3D. Our approach begins by integrating multiview RGB, normal, and coordinate images into a unified representation, termed as 2.5D latents. Next, we adapt pre-trained 2D foundation models for high-fidelity 2.5D generation, utilizing both text and image conditions. Finally, we introduce a lightweight 2.5D-to-3D refiner-decoder framework that efficiently generates detailed 3D representations from 2.5D images. Extensive experiments demonstrate that our model not only excels in generating high-quality 3D objects with coherent structure and color from text and image inputs but also significantly outperforms existing methods in geometry-conditioned texture generation.

3D生成2.5D表示纹理生成联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。