arXiv:2503.13576cs.CV2025-03综述被引 5

梳理文本到图像扩散模型中视觉概念挖掘的四大技术路径

A Comprehensive Survey on Visual Concept Mining in Text-to-image Diffusion Models

  • 按学习、消除、分解、组合四类归纳视觉概念挖掘方法
  • 揭示现有技术在控制力提升上的关键机制与局限性
  • 适合关注可控图像生成与个性化建模的研究者参考

文本到图像扩散模型在从文本提示生成高质量、多样化图像方面取得了显著进展。然而,文本信号的固有局限性常导致模型难以充分捕捉特定概念,从而降低可控性。为此,多种个性化技术通过引入参考图像来挖掘视觉概念表征,以补充文本输入并增强模型可控性。尽管已有进展,但对视觉概念挖掘的系统性探索仍显不足。本文将现有研究归为四个核心方向:概念学习、概念消除、概念分解和概念组合。该分类有助于理解视觉概念挖掘(VCM)技术的基础原理。此外,我们识别出关键挑战,并提出未来研究方向,以推动这一重要且富有前景领域的持续发展。

原文摘要 · Abstract (English)

Text-to-image diffusion models have made significant advancements in generating high-quality, diverse images from text prompts. However, the inherent limitations of textual signals often prevent these models from fully capturing specific concepts, thereby reducing their controllability. To address this issue, several approaches have incorporated personalization techniques, utilizing reference images to mine visual concept representations that complement textual inputs and enhance the controllability of text-to-image diffusion models. Despite these advances, a comprehensive, systematic exploration of visual concept mining remains limited. In this paper, we categorize existing research into four key areas: Concept Learning, Concept Erasing, Concept Decomposition, and Concept Combination. This classification provides valuable insights into the foundational principles of Visual Concept Mining (VCM) techniques. Additionally, we identify key challenges and propose future research directions to propel this important and interesting field forward.

视觉概念挖掘扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。