arXiv:2512.05152cs.CV2025-12被引 1

提升细粒度图像生成质量,解决语义混淆与细节不足问题

EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models

  • 分层嵌入融合上下位类别语义信息,缓解语义混淆
  • 在感知生成阶段引入超分辨率,增强图像细节特征
  • 设计轻量级ProAttention,提升扩散模型效率,适合资源受限场景

扩散模型因其可控性和生成图像的多样性备受关注。然而,基于扩散模型的类别条件生成方法通常聚焦于常见类别,在大规模细粒度图像生成中,仍存在语义信息纠缠和生成图像细节不足的问题。本文提出一种分层嵌入机制,融合父类与子类的语义信息,使扩散模型更好地利用语义内容,缓解语义纠缠。为解决细节不足问题,我们在感知信息生成阶段引入超分辨率技术,通过增强与退化模型提升细粒度图像的细节表现。此外,我们设计了一种高效的ProAttention机制,可有效部署于扩散模型中。在多个公开基准上的大量实验表明,本方法在性能上优于现有最先进的微调方法。

原文摘要 · Abstract (English)

Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models often focus on more common categories. In large-scale fine-grained image generation, issues of semantic information entanglement and insufficient detail in the generated images still persist. This paper attempts to introduce a concept of a tiered embedder in fine-grained image generation, which integrates semantic information from both super and child classes, allowing the diffusion model to better incorporate semantic information and address the issue of semantic entanglement. To address the issue of insufficient detail in fine-grained images, we introduce the concept of super-resolution during the perceptual information generation stage, enhancing the detailed features of fine-grained images through enhancement and degradation models. Furthermore, we propose an efficient ProAttention mechanism that can be effectively implemented in the diffusion model. We evaluate our method through extensive experiments on public benchmarks, demonstrating that our approach outperforms other state-of-the-art fine-tuning methods in terms of performance.

图像生成扩散模型细粒度识别超分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。