arXiv:2512.12800cs.CV2025-12被引 2

从两组图像中分离共性与独特生成因素,提升合成质量。

Learning Common and Salient Generative Factors Between Two Image Datasets

  • 设计新学习策略,区分跨数据集共性与独有特征。
  • 在人脸、动物和医学影像上实现更优的特征分离与生成质量。
  • 无需属性标签,适用于GAN与扩散模型,适合图像生成研究者。

近期图像合成技术实现了高质量图像生成与操控。现有工作主要关注:1)条件化操控(如基于属性修改图像),或2)解耦表征学习(使每个潜在方向对应单一语义属性)。本文聚焦一个较少研究的问题——对比分析(Contrastive Analysis, CA):给定两个图像数据集,目标是分离出两者共享的共性生成因素,以及仅属于其中一个数据集的独特因素。相比依赖属性标签(如眼镜、性别)进行编辑的现有方法,本方法仅使用数据集信号,约束更弱。我们提出一种新颖框架,可适配于GAN与扩散模型,用于学习共性与独特因素。通过设计新的学习策略与损失函数,确保共性与独特因素的有效分离,并保持高质量生成能力。我们在涵盖人脸、动物图像及医学扫描的多样化数据集上进行评估,结果表明该框架在特征分离能力与图像生成质量方面均优于先前方法。

原文摘要 · Abstract (English)

Recent advancements in image synthesis have enabled high-quality image generation and manipulation. Most works focus on: 1) conditional manipulation, where an image is modified conditioned on a given attribute, or 2) disentangled representation learning, where each latent direction should represent a distinct semantic attribute. In this paper, we focus on a different and less studied research problem, called Contrastive Analysis (CA). Given two image datasets, we want to separate the common generative factors, shared across the two datasets, from the salient ones, specific to only one dataset. Compared to existing methods, which use attributes as supervised signals for editing (e.g., glasses, gender), the proposed method is weaker, since it only uses the dataset signal. We propose a novel framework for CA, that can be adapted to both GAN and Diffusion models, to learn both common and salient factors. By defining new and well-adapted learning strategies and losses, we ensure a relevant separation between common and salient factors, preserving a high-quality generation. We evaluate our approach on diverse datasets, covering human faces, animal images and medical scans. Our framework demonstrates superior separation ability and image quality synthesis compared to prior methods.

图像生成对比分析解耦表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。