arXiv:2508.03415cs.CVcs.AI2025-08中稿 · , the copyright wi…被引 1

通过频域引导增强隐空间学习,提升图像翻译的细节与结构一致性。

Learning Latent Representations for Image Translation using Frequency Distributed CycleGAN

  • 引入局部邻域编码与频域监督,捕捉像素级语义与结构信息。
  • 在马到斑马、莫奈风格等数据集上实现更快收敛和更高感知质量。
  • 适合需要高效训练和高保真输出的图像生成任务,如医学影像合成。

本文提出Fd-CycleGAN,一种基于CycleGAN的图像到图像翻译框架,通过增强隐空间表示学习来逼近真实数据分布。该方法结合局部邻域编码(LNE)与频域感知监督,在保留源域结构一致性的同时,捕获细粒度的局部像素语义。采用基于分布的损失函数(包括KL/JS散度与对数相似性度量),显式量化生成图像与真实图像在空间与频域上的分布对齐程度。在Horse2Zebra、Monet2Photo及合成的Strike-off数据集上进行实验,结果表明,相比基线CycleGAN及其他先进方法,本模型在低数据场景下仍具更优的感知质量、更快收敛速度与更强模式多样性。通过有效捕捉局部与全局分布特征,实现更视觉连贯且语义一致的图像转换。结果表明,频域引导的隐空间学习显著提升图像翻译任务的泛化能力,适用于文档修复、艺术风格迁移与医学图像合成。同时对比扩散模型,验证了轻量级对抗式方法在训练效率与输出质量上的优势。

原文摘要 · Abstract (English)

This paper presents Fd-CycleGAN, an image-to-image (I2I) translation framework that enhances latent representation learning to approximate real data distributions. Building upon the foundation of CycleGAN, our approach integrates Local Neighborhood Encoding (LNE) and frequency-aware supervision to capture fine-grained local pixel semantics while preserving structural coherence from the source domain. We employ distribution-based loss metrics, including KL/JS divergence and log-based similarity measures, to explicitly quantify the alignment between real and generated image distributions in both spatial and frequency domains. To validate the efficacy of Fd-CycleGAN, we conduct experiments on diverse datasets -- Horse2Zebra, Monet2Photo, and a synthetically augmented Strike-off dataset. Compared to baseline CycleGAN and other state-of-the-art methods, our approach demonstrates superior perceptual quality, faster convergence, and improved mode diversity, particularly in low-data regimes. By effectively capturing local and global distribution characteristics, Fd-CycleGAN achieves more visually coherent and semantically consistent translations. Our results suggest that frequency-guided latent learning significantly improves generalization in image translation tasks, with promising applications in document restoration, artistic style transfer, and medical image synthesis. We also provide comparative insights with diffusion-based generative models, highlighting the advantages of our lightweight adversarial approach in terms of training efficiency and qualitative output.

图像翻译频域学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。