arXiv:2412.19533cs.CVcs.AI2024-12被引 1

用点标注实现精准主体生成,省去复杂分割标注

P3S-Diffusion:A Selective Subject-driven Generation Framework via Point Supervision

  • 仅需点击参考图中的关键点,自动扩展出主体掩码
  • 多层条件注入+注意力一致性损失,保留主体细粒度特征
  • 适合需要精确控制生成内容的研究者与设计师

当前主体驱动生成研究愈发重视选择性主体特征。然而,在图像中准确选取特定内容(如两只不同狗)仍具挑战。现有方法依赖文本提示或像素掩码,但前者描述精度不足,后者标注成本高。为此,我们提出P3S-Diffusion,一种基于点监督的上下文选择性主体生成框架。该框架仅需少量点标注即可生成主体驱动图像。微调阶段可由点自动生成扩展掩码,无需额外分割模型。掩码用于图像修复与主体表征对齐。通过多层条件注入保留主体精细特征,并引入注意力一致性损失提升训练效果。大量实验表明,该方法在特征保持与图像生成方面表现优异。

原文摘要 · Abstract (English)

Recent research in subject-driven generation increasingly emphasizes the importance of selective subject features. Nevertheless, accurately selecting the content in a given reference image still poses challenges, especially when selecting the similar subjects in an image (e.g., two different dogs). Some methods attempt to use text prompts or pixel masks to isolate specific elements. However, text prompts often fall short in precisely describing specific content, and pixel masks are often expensive. To address this, we introduce P3S-Diffusion, a novel architecture designed for context-selected subject-driven generation via point supervision. P3S-Diffusion leverages minimal cost label (e.g., points) to generate subject-driven images. During fine-tuning, it can generate an expanded base mask from these points, obviating the need for additional segmentation models. The mask is employed for inpainting and aligning with subject representation. The P3S-Diffusion preserves fine features of the subjects through Multi-layers Condition Injection. Enhanced by the Attention Consistency Loss for improved training, extensive experiments demonstrate its excellent feature preservation and image generation capabilities.

图像生成点标注扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。