用多专家模型融合生成图像视频,无需训练即可灵活控制输出。
Product of Experts for Visual Generation
- 通过泊松混合框架在推理时整合多种异构模型知识
- 采用退火重要性采样实现多专家联合分布采样,提升生成质量
- 适合需要灵活控制生成内容的研究者与开发者
现代神经网络在共享数据域(如图像、视频)中捕捉了丰富的先验知识并具备互补性。然而,如何整合来自不同来源的多样化知识——包括视觉生成模型、视觉语言模型,以及人类设计的知识源(如图形引擎和物理模拟器)——仍缺乏探索。我们提出一种训练无关的专家乘积(Product of Experts, PoE)框架,在推理阶段实现异构模型的知识组合。该方法通过退火重要性采样(Annealed Importance Sampling, AIS)从多个专家的联合分布中采样。实验表明,该框架在图像与视频生成任务中表现出色,相比单体模型具有更强的可控性,并支持灵活的用户交互来指定生成目标。
原文摘要 · Abstract (English)
Modern neural models capture rich priors and have complementary knowledge over shared data domains, e.g., images and videos. Integrating diverse knowledge from multiple sources -- including visual generative models, visual language models, and sources with human-crafted knowledge such as graphics engines and physics simulators -- remains under-explored. We propose a Product of Experts (PoE) framework that performs inference-time knowledge composition from heterogeneous models. This training-free approach samples from the product distribution across experts via Annealed Importance Sampling (AIS). Our framework shows practical benefits in image and video synthesis tasks, yielding better controllability than monolithic methods and additionally providing flexible user interfaces for specifying visual generation goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。