arXiv:2411.17784cs.CV2024-11ICCV被引 4

在双曲空间中生成少样本图像,兼顾类别一致性与多样性。

HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation

  • 用双曲空间建模图像层级关系,提升少样本生成质量。
  • 通过调整双曲盘半径,实现对语义多样性的可控调节。
  • 生成过程可解释性强,适合需要精细控制的图像生成任务。

少样本图像生成旨在仅用少量样本生成未见类别的多样且高质量图像。该任务的核心挑战在于平衡类别一致性与图像多样性,二者常相互制约。现有方法对新生成图像属性的控制能力有限。本文提出双曲扩散自编码器(HypDAE),在双曲空间中捕捉已见类别图像间的层级关系。利用预训练基础模型,HypDAE 通过改变随机子码或语义码,为未见类别生成高质量、多样化的图像。更重要的是,双曲表示通过调节双曲盘内的半径,引入了对语义多样性的额外控制。大量实验与可视化表明,HypDAE 在有限数据下显著优于先前方法,实现了类别相关特征保留与图像多样性之间的更优平衡。此外,其生成过程具有高度可控性与可解释性。

原文摘要 · Abstract (English)

Few-shot image generation aims to generate diverse and high-quality images for an unseen class given only a few examples in that class. A key challenge in this task is balancing category consistency and image diversity, which often compete with each other. Moreover, existing methods offer limited control over the attributes of newly generated images. In this work, we propose Hyperbolic Diffusion Autoencoders (HypDAE), a novel approach that operates in hyperbolic space to capture hierarchical relationships among images from seen categories. By leveraging pre-trained foundation models, HypDAE generates diverse new images for unseen categories with exceptional quality by varying stochastic subcodes or semantic codes. Most importantly, the hyperbolic representation introduces an additional degree of control over semantic diversity through the adjustment of radii within the hyperbolic disk. Extensive experiments and visualizations demonstrate that HypDAE significantly outperforms prior methods by achieving a better balance between preserving category-relevant features and promoting image diversity with limited data. Furthermore, HypDAE offers a highly controllable and interpretable generation process.

少样本生成双曲空间图像生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。