30秒生成4K纹理3D模型,解决多视角不一致问题
CaPa: Carve-n-Paint Synthesis for Efficient 4K Textured Mesh Generation
- 分两阶段生成:先用扩散模型建模结构,再用新注意力机制画高分辨率纹理
- 支持最高4K纹理,生成时间小于30秒,几何与纹理一致性好
- 适合游戏、影视等需要快速产出高质量3D资产的工业场景
从文本或视觉输入合成高质量3D资产已成为现代生成建模的核心目标。尽管3D生成算法层出不穷,但仍面临多视角不一致、生成速度慢、保真度低和表面重建等问题。本文提出CaPa——一种切削-绘画框架,可高效生成高保真3D资产。该方法采用两阶段流程,将几何生成与纹理合成解耦:首先利用3D隐空间扩散模型在多视图输入引导下生成几何结构,确保各视角间的一致性;随后通过一种新型、模型无关的时空解耦注意力机制,为给定几何体合成高达4K分辨率的纹理;此外,提出一种3D感知遮挡修复算法,填补未纹理区域,实现整体模型的连贯性。整个流程可在30秒内完成,输出即为可用的商业级3D资产。实验表明,CaPa在纹理保真度与几何稳定性方面均表现优异,树立了实用化、可扩展3D资产生成的新标准。
原文摘要 · Abstract (English)
The synthesis of high-quality 3D assets from textual or visual inputs has become a central objective in modern generative modeling. Despite the proliferation of 3D generation algorithms, they frequently grapple with challenges such as multi-view inconsistency, slow generation times, low fidelity, and surface reconstruction problems. While some studies have addressed some of these issues, a comprehensive solution remains elusive. In this paper, we introduce \textbf{CaPa}, a carve-and-paint framework that generates high-fidelity 3D assets efficiently. CaPa employs a two-stage process, decoupling geometry generation from texture synthesis. Initially, a 3D latent diffusion model generates geometry guided by multi-view inputs, ensuring structural consistency across perspectives. Subsequently, leveraging a novel, model-agnostic Spatially Decoupled Attention, the framework synthesizes high-resolution textures (up to 4K) for a given geometry. Furthermore, we propose a 3D-aware occlusion inpainting algorithm that fills untextured regions, resulting in cohesive results across the entire model. This pipeline generates high-quality 3D assets in less than 30 seconds, providing ready-to-use outputs for commercial applications. Experimental results demonstrate that CaPa excels in both texture fidelity and geometric stability, establishing a new standard for practical, scalable 3D asset generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。