arXiv:2505.20431cs.GRcs.CV2025-05SIGGRAPH被引 6

用文字提示快速生成带细节的3D模型,支持交互式修改。

ART-DECO: Arbitrary Text Guidance for 3D Detailizer Construction

  • 通过文本引导的神经模型,1秒内将粗糙模型转为高精度细节资产。
  • 基于多视角扩散模型蒸馏知识,支持任意结构与风格的一致生成。
  • 适合需要快速创作、风格统一3D内容的设计师和开发者。

我们提出一种3D细节化模型(3D detailizer),可在不到1秒内将粗略的3D形状代理转化为具有精细几何与纹理的高质量资产,仅需输入文本提示作为指导。该模型通过文本定义类别、外观及细粒度风格,而粗略的3D代理则提供结构控制,支持用户编辑调整。重要的是,该细节化模型并非针对单一形状训练,而是从生成模型中蒸馏出通用知识,无需重训即可适用于任意结构的形状生成,且局部细节保持一致风格。训练采用预训练的多视图图像扩散模型(含文本条件),借助分数蒸馏采样(SDS)将知识迁移到细节化模型中,并分两个阶段训练以提升对复杂结构的泛化能力。大量实验表明,相比现有文本到3D模型,本方法在结构可控性下生成质量更优。该模型可在<1秒内完成细化,支持交互式3D创作;用户自定义结构可生成超越主流模型能力的创意型3D资产。我们展示了该方法支持的交互式建模流程及其在风格、结构与物体类别上的强泛化能力。

原文摘要 · Abstract (English)

We introduce a 3D detailizer, a neural model which can instantaneously (in <1s) transform a coarse 3D shape proxy into a high-quality asset with detailed geometry and texture as guided by an input text prompt. Our model is trained using the text prompt, which defines the shape class and characterizes the appearance and fine-grained style of the generated details. The coarse 3D proxy, which can be easily varied and adjusted (e.g., via user editing), provides structure control over the final shape. Importantly, our detailizer is not optimized for a single shape; it is the result of distilling a generative model, so that it can be reused, without retraining, to generate any number of shapes, with varied structures, whose local details all share a consistent style and appearance. Our detailizer training utilizes a pretrained multi-view image diffusion model, with text conditioning, to distill the foundational knowledge therein into our detailizer via Score Distillation Sampling (SDS). To improve SDS and enable our detailizer architecture to learn generalizable features over complex structures, we train our model in two training stages to generate shapes with increasing structural complexity. Through extensive experiments, we show that our method generates shapes of superior quality and details compared to existing text-to-3D models under varied structure control. Our detailizer can refine a coarse shape in less than a second, making it possible to interactively author and adjust 3D shapes. Furthermore, the user-imposed structure control can lead to creative, and hence out-of-distribution, 3D asset generations that are beyond the current capabilities of leading text-to-3D generative models. We demonstrate an interactive 3D modeling workflow our method enables, and its strong generalizability over styles, structures, and object categories.

3D生成文本引导实时渲染细节增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。