解决多数据集合并时条件缺失问题,实现精准可控生成。
Diffusion Models with Double Guidance: Generate with aggregated datasets
- 引入双引导机制,利用多个数据集中的条件信息联合生成
- 在分子和图像生成任务中,条件匹配度与可控性均优于基线
- 适用于无完整标注样本的复杂条件生成场景
为训练高性能生成模型构建大规模数据集通常成本高昂,尤其当需提供属性或标注时。因此,合并现有数据集成为常见策略。然而,不同数据集的属性集合常不一致,简单拼接会导致块状条件缺失,给联合使用多个属性作为条件的条件生成带来挑战,限制了模型的可控性与适用性。为此,我们提出一种新生成方法——双引导扩散模型,可在无训练样本同时包含所有条件的情况下,仍实现精确的条件生成。该方法无需联合标注即可严格控制多重条件。我们在分子和图像生成任务中验证其有效性,结果表明其在目标条件分布对齐和缺失条件下可控性方面均优于现有基线。
原文摘要 · Abstract (English)
Creating large-scale datasets for training high-performance generative models is often prohibitively expensive, especially when associated attributes or annotations must be provided. As a result, merging existing datasets has become a common strategy. However, the sets of attributes across datasets are often inconsistent, and their naive concatenation typically leads to block-wise missing conditions. This presents a significant challenge for conditional generative modeling when the multiple attributes are used jointly as conditions, thereby limiting the model's controllability and applicability. To address this issue, we propose a novel generative approach, Diffusion Model with Double Guidance, which enables precise conditional generation even when no training samples contain all conditions simultaneously. Our method maintains rigorous control over multiple conditions without requiring joint annotations. We demonstrate its effectiveness in molecular and image generation tasks, where it outperforms existing baselines both in alignment with target conditional distributions and in controllability under missing condition settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。