用扩散模型最大化数据流形熵,实现无不确定性依赖的创新探索。
Provable Maximum Entropy Manifold Exploration via Diffusion Models
- 将探索任务转化为预训练扩散模型隐式定义流形上的熵最大化。
- 通过分数函数关联熵与密度估计,实现可扩展的探索算法。
- 适用于科学发现等需生成新设计的场景,无需显式不确定性计算。
探索在解决现实世界决策问题(如科学发现)中至关重要,目标是生成真正新颖的设计而非模仿现有数据分布。本文提出一种新框架,将探索视为由预训练扩散模型隐式定义的数据流形上的熵最大化问题。针对密度估计这一实践中的难题,我们揭示了扩散模型诱导密度的熵与其得分函数之间的基本联系,从而构建基于镜面下降的算法,通过逐步微调预训练扩散模型来解决探索问题。在合理假设下,我们证明该算法收敛至最优探索性扩散模型,利用了镜面流的最新理解。我们在合成数据和高维文本到图像扩散模型上进行了实证评估,结果表明方法具有潜力。
原文摘要 · Abstract (English)
Exploration is critical for solving real-world decision-making problems such as scientific discovery, where the objective is to generate truly novel designs rather than mimic existing data distributions. In this work, we address the challenge of leveraging the representational power of generative models for exploration without relying on explicit uncertainty quantification. We introduce a novel framework that casts exploration as entropy maximization over the approximate data manifold implicitly defined by a pre-trained diffusion model. Then, we present a novel principle for exploration based on density estimation, a problem well-known to be challenging in practice. To overcome this issue and render this method truly scalable, we leverage a fundamental connection between the entropy of the density induced by a diffusion model and its score function. Building on this, we develop an algorithm based on mirror descent that solves the exploration problem as sequential fine-tuning of a pre-trained diffusion model. We prove its convergence to the optimal exploratory diffusion model under realistic assumptions by leveraging recent understanding of mirror flows. Finally, we empirically evaluate our approach on both synthetic and high-dimensional text-to-image diffusion, demonstrating promising results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。