arXiv:2603.20725cs.CV2026-03被引 1

用可学习的用户嵌入实现精准图像生成个性化。

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

  • 为每个用户建模可学习偏好嵌入,融合文本提示生成图像。
  • 在相同历史长度下,偏好对齐度与文本一致性显著提升。
  • 适合需要个性化图像生成的场景,尤其用户数据稀少时表现佳。

文本到图像生成虽快速进步,但仍难以捕捉细微用户偏好。现有方法多依赖多模态大语言模型推断偏好,但生成的提示或潜在码常不准确,导致个性化效果不佳。本文提出Premier,一种新的偏好调制框架。它将每位用户的偏好表示为可学习嵌入,并引入偏好适配器将用户嵌入与文本提示融合;进一步利用融合后的偏好嵌入调制生成过程,实现细粒度控制。为增强个体偏好差异并提升输出与用户风格的一致性,引入分散损失,强制用户嵌入间分离。当用户数据稀缺时,新用户通过已有训练嵌入的线性组合表示,实现有效泛化。实验表明,Premier在相同历史长度下优于先前方法,在偏好对齐、文本一致性、ViPer代理指标及专家评估上均表现更优。

原文摘要 · Abstract (English)

Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language models to infer user preferences, but the derived prompts or latent codes rarely reflect them faithfully, leading to suboptimal personalization. We present Premier, a novel preference modulation framework for personalized image generation. Premier represents each user's preference as a learnable embedding and introduces a preference adapter that fuses the user embedding with the text prompt. To enable accurate and fine-grained preference control, the fused preference embedding is further used to modulate the generative process. To enhance the distinctness of individual preference and improve alignment between outputs and user-specific styles, we incorporate a dispersion loss that enforces separation among user embeddings. When user data are scarce, new users are represented as linear combinations of existing preference embeddings learned during training, enabling effective generalization. Experiments show that Premier outperforms prior methods under the same history length, achieving stronger preference alignment and superior performance on text consistency, ViPer proxy metrics, and expert evaluations.

图像生成个性化用户嵌入偏好建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。