用模糊规则引导多路径扩散,提升复杂图像生成稳定性与质量
Diffusion Fuzzy System: Fuzzy Rule Guided Latent Multi-Path Diffusion Modeling
- 多路径分别学习不同图像特征,由模糊规则动态协调
- 在三个数据集上训练更稳定,收敛更快,生成图像更逼真
- 适合需要高保真、多特征一致生成的场景
扩散模型因其生成高分辨率真实图像的能力而成为主流技术。然而,面对特征差异显著的图像集合时,传统方法难以捕捉复杂特征并易产生矛盾结果。现有方案通过多路径学习图像不同区域并融合,但存在路径间协调低效和计算成本高的问题。本文提出扩散模糊系统(DFS),一种基于模糊规则引导的潜在空间多路径扩散模型。DFS将各路径专用于特定类型图像特征的学习,突破了多路径模型对异质特征的捕捉瓶颈;采用基于规则链的推理机制,动态引导扩散过程,实现路径间高效协同;引入基于模糊隶属度的潜在空间压缩机制,显著降低多路径扩散的计算开销。在LSUN Bedroom、LSUN Church和MS COCO三个公开数据集上的实验表明,DFS实现了更稳定的训练过程和更快的收敛速度,优于现有单路径与多路径扩散模型。同时,在图像质量、文本-图像一致性以及生成图像与目标参考之间的匹配精度方面均表现更优。
原文摘要 · Abstract (English)
Diffusion models have emerged as a leading technique for generating images due to their ability to create high-resolution and realistic images. Despite their strong performance, diffusion models still struggle in managing image collections with significant feature differences. They often fail to capture complex features and produce conflicting results. Research has attempted to address this issue by learning different regions of an image through multiple diffusion paths and then combining them. However, this approach leads to inefficient coordination among multiple paths and high computational costs. To tackle these issues, this paper presents a Diffusion Fuzzy System (DFS), a latent-space multi-path diffusion model guided by fuzzy rules. DFS offers several advantages. First, unlike traditional multi-path diffusion methods, DFS uses multiple diffusion paths, each dedicated to learning a specific class of image features. By assigning each path to a different feature type, DFS overcomes the limitations of multi-path models in capturing heterogeneous image features. Second, DFS employs rule-chain-based reasoning to dynamically steer the diffusion process and enable efficient coordination among multiple paths. Finally, DFS introduces a fuzzy membership-based latent-space compression mechanism to reduce the computational costs of multi-path diffusion effectively. We tested our method on three public datasets: LSUN Bedroom, LSUN Church, and MS COCO. The results show that DFS achieves more stable training and faster convergence than existing single-path and multi-path diffusion models. Additionally, DFS surpasses baseline models in both image quality and alignment between text and images, and also shows improved accuracy when comparing generated images to target references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。