通过几何映射在隐空间对齐跨域特征,避免生成模型的模式崩溃问题。
GMapLatent: Geometric Mapping in Latent Space
- 用重心变换、最优传输和调和映射构建规范隐空间,实现严格对应
- 在灰度与彩色图像上验证,生成质量优于现有方法
- 适合需要高精度跨域生成的研究者或工业应用
基于编码器-解码器架构的跨域生成模型在生成真实图像方面备受关注,其中域对齐是决定生成精度的关键。传统对齐方法直接处理初始分布,但不匹配或混合的簇可能导致解码器出现模式崩溃和混合问题,损害模型泛化能力。本文提出GMapLatent模型,通过几何映射在规范隐空间中精确对齐跨域隐空间,避免上述问题。核心方法包括:(1)通过重心平移、最优传输融合和约束调和映射将隐空间转换至规范参数域;(2)在规范参数域上计算带簇约束的几何配准。该过程实现了新隐空间间的双射映射,并精准对齐簇对。跨域生成通过嵌入对齐隐空间的编码器-解码器流程实现。在灰度与彩色图像上的实验验证了GMapLatent的高效性、有效性与适用性,其性能显著优于现有模型。
原文摘要 · Abstract (English)
Cross-domain generative models based on encoder-decoder AI architectures have attracted much attention in generating realistic images, where domain alignment is crucial for generation accuracy. Domain alignment methods usually deal directly with the initial distribution; however, mismatched or mixed clusters can lead to mode collapse and mixture problems in the decoder, compromising model generalization capabilities. In this work, we innovate a cross-domain alignment and generation model that introduces a canonical latent space representation based on geometric mapping to align the cross-domain latent spaces in a rigorous and precise manner, thus avoiding mode collapse and mixture in the encoder-decoder generation architectures. We name this model GMapLatent. The core of the method is to seamlessly align latent spaces with strict cluster correspondence constraints using the canonical parameterizations of cluster-decorated latent spaces. We first (1) transform the latent space to a canonical parameter domain by composing barycenter translation, optimal transport merging and constrained harmonic mapping, and then (2) compute geometric registration with cluster constraints over the canonical parameter domains. This process realizes a bijective (one-to-one and onto) mapping between newly transformed latent spaces and generates a precise alignment of cluster pairs. Cross-domain generation is then achieved through the aligned latent spaces embedded in the encoder-decoder pipeline. Experiments on gray-scale and color images validate the efficiency, efficacy and applicability of GMapLatent, and demonstrate that the proposed model has superior performance over existing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。