arXiv:2412.16156cs.CVcs.LG2024-12被引 10

用个性化生成图像提升小样本视觉任务的表示学习

Personalized Representation from Personalized Generation

  • 利用扩散模型从少量真实图生成个性化数据,用于对比学习
  • 在识别、分割等任务上显著提升小样本个性化表示性能
  • 适合研究个性化视觉表征与生成式学习的学者

当前视觉模型在通用下游任务中表现优异,但对细粒度、数据稀缺的个性化视觉任务仍不明确。近期工作已成功将合成数据用于通用表征学习,而文本到图像扩散模型仅需少数真实样本即可生成个性化图像。本文探索二者结合的可能性,形式化了使用个性化合成数据学习个性化表征的挑战——该表征需编码目标对象知识,并可灵活应用于相关下游任务。我们构建了一个评估套件,包含两个现有数据集的重构版本及一个专为此任务设计的新数据集,并提出一种创造性地利用图像生成器的对比学习方法。实验表明,该方法在多样化的下游任务(如识别、分割)中均显著提升个性化表征学习效果,并分析了影响性能的关键生成策略。

原文摘要 · Abstract (English)

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully applied synthetic data to general-purpose representation learning, while advances in T2I diffusion models have enabled the generation of personalized images from just a few real examples. Here, we explore a potential connection between these ideas, and formalize the challenge of using personalized synthetic data to learn personalized representations, which encode knowledge about an object of interest and may be flexibly applied to any downstream task relating to the target object. We introduce an evaluation suite for this challenge, including reformulations of two existing datasets and a novel dataset explicitly constructed for this purpose, and propose a contrastive learning approach that makes creative use of image generators. We show that our method improves personalized representation learning for diverse downstream tasks, from recognition to segmentation, and analyze characteristics of image generation approaches that are key to this gain.

个性化表征生成模型小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。