用扩散模型实现语义可控的反事实图像生成,兼顾身份一致与因果真实性。
Diffusion Counterfactual Generation with Semantic Abduction
- 基于皮尔利因果理论,通过语义抽象实现反事实编辑
- 首次在扩散模型中实现高阶语义身份保留,提升生成一致性
- 适合关注可解释生成与因果推理的研究者
反事实图像生成面临身份保持、感知质量与因果模型忠实度等挑战。现有自编码框架虽具备可操纵的语义潜在空间,但存在可扩展性差和保真度不足的问题。扩散模型在视觉质量、人类感知对齐和表征学习方面表现优异,为改进反事实图像编辑提供了新可能。本文提出一套基于扩散模型的因果机制,引入空间、语义与动态抽象概念,构建一个将语义表示融入扩散模型的通用框架,通过反事实推理实现图像编辑。据我们所知,这是首个考虑扩散模型中高层语义身份保留的工作,并展示了语义控制如何在因果忠实性与身份保真之间实现合理权衡。
原文摘要 · Abstract (English)
Counterfactual image generation presents significant challenges, including preserving identity, maintaining perceptual quality, and ensuring faithfulness to an underlying causal model. While existing auto-encoding frameworks admit semantic latent spaces which can be manipulated for causal control, they struggle with scalability and fidelity. Advancements in diffusion models present opportunities for improving counterfactual image editing, having demonstrated state-of-the-art visual quality, human-aligned perception and representation learning capabilities. Here, we present a suite of diffusion-based causal mechanisms, introducing the notions of spatial, semantic and dynamic abduction. We propose a general framework that integrates semantic representations into diffusion models through the lens of Pearlian causality to edit images via a counterfactual reasoning process. To our knowledge, this is the first work to consider high-level semantic identity preservation for diffusion counterfactuals and to demonstrate how semantic control enables principled trade-offs between faithful causal control and identity preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。