发现判别模型暗藏生成能力,无需训练即可高质量合成图像
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
- 通过多尺度优化释放CLIP的隐含生成能力
- 生成图像保持自然统计特性,避免对抗性伪影
- 适用于文生图、风格迁移等场景,适合研究模型本质的学者
我们证明判别模型本质上具备强大的生成能力,挑战了判别与生成架构的根本区别。所提方法Direct Ascent Synthesis(DAS)通过多分辨率优化CLIP模型表征,揭示其潜在生成能力。传统反演常产生对抗性模式,而DAS在1×1至224×224多尺度上分解优化,无需额外训练即可实现高质量图像合成。该方法不仅支持文本到图像生成、风格迁移等多样化应用,还保持自然图像统计特性($1/f^2$频谱),引导生成避开非鲁棒的对抗性模式。结果表明,标准判别模型蕴含远超以往认知的生成知识,为模型可解释性及对抗样本与自然图像合成的关系提供了新视角。
原文摘要 · Abstract (English)
We demonstrate that discriminative models inherently contain powerful generative capabilities, challenging the fundamental distinction between discriminative and generative architectures. Our method, Direct Ascent Synthesis (DAS), reveals these latent capabilities through multi-resolution optimization of CLIP model representations. While traditional inversion attempts produce adversarial patterns, DAS achieves high-quality image synthesis by decomposing optimization across multiple spatial scales (1x1 to 224x224), requiring no additional training. This approach not only enables diverse applications -- from text-to-image generation to style transfer -- but maintains natural image statistics ($1/f^2$ spectrum) and guides the generation away from non-robust adversarial patterns. Our results demonstrate that standard discriminative models encode substantially richer generative knowledge than previously recognized, providing new perspectives on model interpretability and the relationship between adversarial examples and natural image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。