用视频扩散先验提升生成模型输出质量,解决多模型流程中的分布不匹配问题。
Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier
- 利用视频扩散先验引导前序模型去噪,对齐下游需求。
- 在视频与3D生成任务中均显著提升质量,通用性强。
- 可插件式部署,适合复杂生成流程或独立优化单个模型。
在实际AI工作流中,复杂任务常涉及多个生成模型的串联,例如在2D图像生成后接视频或3D生成模型。然而,上游模型输出与下游模型期望输入之间的分布差异,常导致整体生成质量下降。为此,我们提出Uni-Classifier(Uni-C),一个简单高效的即插即用模块,利用视频扩散先验引导前序模型的去噪过程,使其输出更符合下游要求。Uni-C亦可独立使用,以提升单一生成模型的输出质量。大量实验表明,无论在流程化设置还是独立使用场景下,Uni-C均能持续提升视频与3D生成质量,展现出优异的泛化能力与通用性。
原文摘要 · Abstract (English)
In practical AI workflows, complex tasks often involve chaining multiple generative models, such as using a video or 3D generation model after a 2D image generator. However, distributional mismatches between the output of upstream models and the expected input of downstream models frequently degrade overall generation quality. To address this issue, we propose Uni-Classifier (Uni-C), a simple yet effective plug-and-play module that leverages video diffusion priors to guide the denoising process of preceding models, thereby aligning their outputs with downstream requirements. Uni-C can also be applied independently to enhance the output quality of individual generative models. Extensive experiments across video and 3D generation tasks demonstrate that Uni-C consistently improves generation quality in both workflow-based and standalone settings, highlighting its versatility and strong generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。