arXiv:2504.12844cs.CV2025-04IJCV被引 8

用多模态信息引导生成,让修复图像更真实且保持原图一致。

High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion

  • 融合边缘、语义等多模态信息增强结构感知
  • 提出F&W+潜空间解决颜色与语义不一致问题
  • 适合处理大范围破损的高保真图像修复

生成对抗网络(GAN)反演在图像修复中表现优异,能利用未遮挡区域恢复缺失纹理。但现有方法忽视输入与输出未遮挡区域必须一致这一硬约束,导致性能下降。同时,多数方法仅依赖单一图像模态,忽略辅助信息。为此,我们提出新方法MMInvertFill,包含多模态引导编码器和带F&W+潜空间的GAN生成器。编码器通过门控掩码感知注意力模块,融合边缘、语义分割等多尺度结构信息;预调制模块将结构编码为风格向量。F&W+潜空间有效缓解颜色偏差与语义不一致问题。此外,引入简单高效的Soft-update Mean Latent模块,捕捉多样化域内模式,生成高质量纹理。在六个挑战性数据集上的实验表明,该方法在定性和定量上均优于现有最优方法,并可有效处理域外图像修复。

原文摘要 · Abstract (English)

Generative Adversarial Network (GAN) inversion have demonstrated excellent performance in image inpainting that aims to restore lost or damaged image texture using its unmasked content. Previous GAN inversion-based methods usually utilize well-trained GAN models as effective priors to generate the realistic regions for missing holes. Despite excellence, they ignore a hard constraint that the unmasked regions in the input and the output should be the same, resulting in a gap between GAN inversion and image inpainting and thus degrading the performance. Besides, existing GAN inversion approaches often consider a single modality of the input image, neglecting other auxiliary cues in images for improvements. Addressing these problems, we propose a novel GAN inversion approach, dubbed MMInvertFill, for image inpainting. MMInvertFill contains primarily a multimodal guided encoder with a pre-modulation and a GAN generator with F&W+ latent space. Specifically, the multimodal encoder aims to enhance the multi-scale structures with additional semantic segmentation edge texture modalities through a gated mask-aware attention module. Afterwards, a pre-modulation is presented to encode these structures into style vectors. To mitigate issues of conspicuous color discrepancy and semantic inconsistency, we introduce the F&W+ latent space to bridge the gap between GAN inversion and image inpainting. Furthermore, in order to reconstruct faithful and photorealistic images, we devise a simple yet effective Soft-update Mean Latent module to capture more diversified in-domain patterns for generating high-fidelity textures for massive corruptions. In our extensive experiments on six challenging datasets, we show that our MMInvertFill qualitatively and quantitatively outperforms other state-of-the-arts and it supports the completion of out-of-domain images effectively.

图像修复GAN反演多模态高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。