让扩散模型准确生成任意逻辑组合的属性数据,解决传统方法偏差问题。
CoInD: Enabling Logical Compositions in Diffusion Models

- 通过最小化费雪散度强制条件边缘分布独立,实现逻辑组合生成。
- 在包含NOT操作或部分组合训练时,生成质量显著优于现有方法。
- 适合需要精确控制属性组合生成的研究者和开发者使用。
如何让生成模型采样出具有任意逻辑组合的统计独立属性数据?当前主流方法假设各属性条件边缘分布独立,基于其组合分布进行采样。本文指出,标准条件扩散模型即使在训练中观察到所有组合,仍违反该独立性假设,且当仅观察部分组合时,偏差更为严重。为此提出CoInD,通过最小化联合分布与边缘分布间的费雪散度,显式强制条件边缘分布的统计独立性。理论优势在定性和定量实验中均体现:对任意逻辑组合的生成更忠实、更可控。尤其在现有依赖独立性假设的方法难以处理的场景下——如涉及NOT操作或仅部分组合可用训练时——效果更显著。
原文摘要 · Abstract (English)
How can we learn generative models to sample data with arbitrary logical compositions of statistically independent attributes? The prevailing solution is to sample from distributions expressed as a composition of attributes' conditional marginal distributions under the assumption that they are statistically independent. This paper shows that standard conditional diffusion models violate this assumption, even when all attribute compositions are observed during training. And, this violation is significantly more severe when only a subset of the compositions is observed. We propose CoInD to address this problem. It explicitly enforces statistical independence between the conditional marginal distributions by minimizing Fisher's divergence between the joint and marginal distributions. The theoretical advantages of CoInD are reflected in both qualitative and quantitative experiments, demonstrating a significantly more faithful and controlled generation of samples for arbitrary logical compositions of attributes. The benefit is more pronounced for scenarios that current solutions relying on the assumption of conditionally independent marginals struggle with, namely, logical compositions involving the NOT operation and when only a subset of compositions are observed during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。