提出新型判别器,同时提升生成质量与条件一致性
SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- 设计三合一判别器,分别评估真实度、条件匹配度和动态平衡目标
- 在图像分类条件下生成质量超越现有最佳方法
- 适用于文本到图像生成,对条件不匹配更敏感
深度生成模型在生成复杂内容方面取得显著进展,但条件生成仍是核心挑战。现有条件生成对抗网络的判别器难以兼顾样本真实性和条件一致性。为此,我们提出一种新型判别器设计,集成无条件判别、匹配感知监督和自适应加权机制。具体地,引入自然度与对齐度之和(SONA),在最后一层采用独立投影分别建模自然度(真实性)与对齐度,并通过专用目标函数与自适应权重机制实现优化。大量实验表明,在类别条件生成任务中,该方法在样本质量和条件对齐性上均优于当前最优方法。此外,其在文本到图像生成中也表现出色,验证了方法的通用性与鲁棒性。
原文摘要 · Abstract (English)
Deep generative models have made significant advances in generating complex content, yet conditional generation remains a fundamental challenge. Existing conditional generative adversarial networks often struggle to balance the dual objectives of assessing authenticity and conditional alignment of input samples within their conditional discriminators. To address this, we propose a novel discriminator design that integrates three key capabilities: unconditional discrimination, matching-aware supervision to enhance alignment sensitivity, and adaptive weighting to dynamically balance all objectives. Specifically, we introduce Sum of Naturalness and Alignment (SONA), which employs separate projections for naturalness (authenticity) and alignment in the final layer with an inductive bias, supported by dedicated objective functions and an adaptive weighting mechanism. Extensive experiments on class-conditional generation tasks show that \ours achieves superior sample quality and conditional alignment compared to state-of-the-art methods. Furthermore, we demonstrate its effectiveness in text-to-image generation, confirming the versatility and robustness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。