用2D扩散模型指导3D风格迁移,让生成更灵活多样。
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion

- 用预训练2D扩散模型提供通用风格先验,指导3D隐变量优化。
- 在多种3D生成模型上实现跨分布风格生成,效果稳定且可即插即用。
- 适合需要灵活风格控制的3D内容创作场景,如游戏与虚拟现实。
3D资产生成在游戏和虚拟现实等领域至关重要,能从单张或多张图像快速合成高保真3D物体。在此基础上,实现可控风格生成成为重要方向。然而,现有方法通常依赖于与训练数据分布相似的风格图像,面对分布外(OOD)风格时性能显著下降甚至失败。为此,我们提出DiLAST:基于2D扩散模型的3D风格迁移隐变量唤醒方法。通过利用预训练2D扩散模型作为教师,提供丰富的通用风格先验,结合扩散引导对渲染视图进行风格对齐,优化结构化3D隐变量以实现风格化。我们发现该限制并非源于模型容量不足,而是结构化3D隐变量未被充分使用,其本身具有强表达能力。即使训练数据有限,3D生成模型也能借助2D扩散引导,在隐空间中沿特定方向去噪,从而生成多样化、分布外风格。在多种数据集和多个3D生成骨干网络上的实验表明,该方法有效且具备即插即用特性。
原文摘要 · Abstract (English)
3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a single or multiple images. Building on this capability, enabling style-controllable generation naturally emerges as an important and desirable direction. However, existing approaches typically rely on style images that lie within or are similar to the training distribution of 3D generation models. When presented with out-of-distribution (OOD) styles, their performance degrades significantly or even fails. To address this limitation, we introduce \textbf{DiLAST}: 2D Diffusion-based Latent Awakening for 3D Style Transfer. Specifically, we leverage a pretrained 2D diffusion model as a teacher to provide rich and generalizable style priors. By aligning rendered views with the target style under diffusion-based guidance, our method optimizes the structured 3D latent representations for stylization. We observe that this limitation stems not from insufficient model capacity, but from the underutilization of structured 3D latents, which are inherently expressive. Despite being trained on comparatively limited data, 3D generation models can leverage 2D diffusion guidance to steer denoising toward specific directions in latent space, thereby producing diverse, OOD styles. Extensive experiments across diverse data and multiple 3D generation backbones demonstrate the effectiveness and plug-and-play nature of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。