用薛定谔桥直接生成3D,减少伪影,提升质量。
Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D Generation
- 将3D生成建模为从当前渲染到目标分布的最优轨迹。
- 在多个数据集上,生成质量超越现有方法,且使用更低的CFG值。
- 适合关注高质量3D生成与扩散模型优化的研究者。
基于优化的文本到3D生成方法通常依赖于从预训练文本到图像扩散模型中蒸馏知识,如得分蒸馏采样(SDS),但常引入过饱和和过度平滑等伪影。本文提出将生成过程建模为从当前渲染分布到目标分布的最优直接传输轨迹,从而实现高质量生成并降低分类器无关引导(CFG)值。理论上,我们证明了SDS是薛定谔桥框架的简化实例,其逆过程在特定条件下(如一端为高斯噪声)退化为预训练扩散模型的得分函数。基于此,我们提出轨迹中心蒸馏(TraCe),一个新型文本到3D生成框架,将薛定谔桥的可追踪数学结构重构为从当前渲染到文本条件去噪目标的扩散桥,并训练一个LoRA适配模型以学习该轨迹的得分动态,实现鲁棒的3D优化。大量实验表明,TraCe在多个基准上持续优于当前最优技术。
原文摘要 · Abstract (English)
Recent advancements in optimization-based text-to-3D generation heavily rely on distilling knowledge from pre-trained text-to-image diffusion models using techniques like Score Distillation Sampling (SDS), which often introduce artifacts such as over-saturation and over-smoothing into the generated 3D assets. In this paper, we address this essential problem by formulating the generation process as learning an optimal, direct transport trajectory between the distribution of the current rendering and the desired target distribution, thereby enabling high-quality generation with smaller Classifier-free Guidance (CFG) values. At first, we theoretically establish SDS as a simplified instance of the Schrödinger Bridge framework. We prove that SDS employs the reverse process of an Schrödinger Bridge, which, under specific conditions (e.g., a Gaussian noise as one end), collapses to SDS's score function of the pre-trained diffusion model. Based upon this, we introduce Trajectory-Centric Distillation (TraCe), a novel text-to-3D generation framework, which reformulates the mathematically trackable framework of Schrödinger Bridge to explicitly construct a diffusion bridge from the current rendering to its text-conditioned, denoised target, and trains a LoRA-adapted model on this trajectory's score dynamics for robust 3D optimization. Comprehensive experiments demonstrate that TraCe consistently achieves superior quality and fidelity to state-of-the-art techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。