arXiv:2508.03481cs.CVcs.AI2025-08ICCV被引 6

通过条件级建模实现文本到图像生成的个性化,无需大量调参

Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models

论文配图:Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
图 1 · 摘自论文原文
  • 在潜在空间中用适配器融合用户画像进行个性化控制
  • 在大规模数据集上表现优异,无需额外微调
  • 适合希望快速定制生成内容的开发者与设计师

文本到图像扩散模型中的个性化生成旨在以最少用户干预自然融入个人偏好。然而,现有方法主要依赖提示层建模与大模型,常因文本编码器输入令牌容量有限导致个性化不准确。为此,我们提出DrUM,一种将用户画像与基于Transformer的适配器结合的新方法,实现潜在空间中的条件级建模。DrUM在大规模数据集上表现强劲,可无缝集成开源文本编码器,兼容主流基础文本到图像模型,无需额外微调。

原文摘要 · Abstract (English)

Personalized generation in T2I diffusion models aims to naturally incorporate individual user preferences into the generation process with minimal user intervention. However, existing studies primarily rely on prompt-level modeling with large-scale models, often leading to inaccurate personalization due to the limited input token capacity of T2I diffusion models. To address these limitations, we propose DrUM, a novel method that integrates user profiling with a transformer-based adapter to enable personalized generation through condition-level modeling in the latent space. DrUM demonstrates strong performance on large-scale datasets and seamlessly integrates with open-source text encoders, making it compatible with widely used foundation T2I models without requiring additional fine-tuning.

个性化生成扩散模型条件建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。