arXiv:2601.11444cs.LGcs.CV2026-01中稿 · Transactions on Ma…被引 1

探究扩散模型集成能否提升生成质量,发现数学性能好但视觉效果不升反降。

When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models

  • 用多种聚合方式集成扩散模型得分函数,提升分数匹配与似然。
  • 在CIFAR-10和FFHQ上,集成后FID指标未改善,视觉质量无提升。
  • 理论分析揭示得分相加机制,为引导生成等技术提供新视角。

扩散模型如今能生成高质量、多样化的样本,研究重点逐渐转向更强模型。尽管集成是提升监督模型的成熟方法,但在无条件基于得分的扩散模型中仍鲜有探索。本文研究集成是否对生成建模带来实际收益。结果表明,虽然集成得分通常能降低分数匹配损失并提高模型似然,却未能持续提升图像数据集上的感知质量指标(如FID)。该现象在多种聚合规则下于CIFAR-10和FFHQ上均得到验证,涵盖深度集成、蒙特卡洛丢弃。我们尝试通过分析得分估计与图像质量的关系解释这一差异。此外,我们在表格数据上使用随机森林也发现某种聚合策略更优。最后,我们提供了关于得分模型相加的理论见解,不仅揭示了集成机制,也为引导生成等模型组合技术提供了支持。

原文摘要 · Abstract (English)

Diffusion models now generate high-quality, diverse samples, with an increasing focus on more powerful models. Although ensembling is a well-known way to improve supervised models, its application to unconditional score-based diffusion models remains largely unexplored. In this work we investigate whether it provides tangible benefits for generative modelling. We find that while ensembling the scores generally improves the score-matching loss and model likelihood, it fails to consistently enhance perceptual quality metrics such as FID on image datasets. We confirm this observation across a breadth of aggregation rules using Deep Ensembles, Monte Carlo Dropout, on CIFAR-10 and FFHQ. We attempt to explain this discrepancy by investigating possible explanations, such as the link between score estimation and image quality. We also look into tabular data through random forests, and find that one aggregation strategy outperforms the others. Finally, we provide theoretical insights into the summing of score models, which shed light not only on ensembling but also on several model composition techniques (e.g. guidance).

扩散模型模型集成生成质量理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。