arXiv:2410.08207cs.CVcs.LG2024-10中稿 · CVPR被引 3

让离散扩散模型实现精准内容编辑,无需预设掩码

DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models

  • 记录反向扩散过程的噪声与掩码序列,实现离散数据逆推
  • 在图像和文本任务中保持高保真度,支持灵活编辑
  • 适用于VQ-Diffusion、Paella等模型,适合精细内容控制场景

离散扩散模型在图像生成和掩码语言建模等任务中表现优异,但在可控内容编辑方面存在局限。本文提出DICE(离散逆推用于可控编辑),首个实现离散扩散模型精确逆推的方法,涵盖多项式扩散和掩码生成模型。通过记录反向扩散过程中的噪声序列和掩码模式,DICE可在不依赖预定义掩码或注意力操作的情况下,实现离散数据的准确重建与灵活编辑。我们在VQ-Diffusion、Paella和RoBERTa等模型上验证了DICE的有效性,结果表明其在保持高数据保真度的同时显著提升了编辑能力,为离散空间中的细粒度内容操控提供了新可能。

原文摘要 · Abstract (English)

Discrete diffusion models have achieved success in tasks like image generation and masked language modeling but face limitations in controlled content editing. We introduce DICE (Discrete Inversion for Controllable Editing), the first approach to enable precise inversion for discrete diffusion models, including multinomial diffusion and masked generative models. By recording noise sequences and masking patterns during the reverse diffusion process, DICE enables accurate reconstruction and flexible editing of discrete data without the need for predefined masks or attention manipulation. We demonstrate the effectiveness of DICE across both image and text domains, evaluating it on models such as VQ-Diffusion, Paella, and RoBERTa. Our results show that DICE preserves high data fidelity while enhancing editing capabilities, offering new opportunities for fine-grained content manipulation in discrete spaces.

扩散模型内容编辑离散生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。