arXiv:2512.06020cs.CVcs.AI2025-12被引 4

用多模态模型精准捕捉用户审美偏好,生成更符合个人口味的图像。

PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation

  • 利用多模态大模型提取用户视觉偏好特征,结合问答任务捕捉细粒度语义。
  • 通过跨用户与同用户对比任务,分离出真正反映偏好的关键特征。
  • 设计对齐损失使文本编码器兼容多模态嵌入,提升个性化生成效果。

偏好条件化图像生成旨在让生成模型适应个体用户,产出不仅符合文本提示、更能体现个人审美的图像。现有方法或难以捕捉细微偏好,或缺乏有效编码个性化视觉信号的机制。本文提出一种多模态框架,利用多模态大语言模型(MLLM)提取丰富用户表征,并注入基于扩散的图像生成过程。通过面向偏好的视觉问答任务训练MLLM以捕捉细粒度语义线索。为分离偏好相关特征,引入两项互补探针任务:跨用户区分任务用于区分不同用户,同用户区分任务用于分离喜欢与不喜欢内容。为确保与扩散文本编码器兼容,设计基于最大均值差异的对齐损失,在保留多模态结构的同时弥合模态差距。最终嵌入用于条件化生成器,实现对提示与用户偏好的双重忠实遵循。大量实验表明,该方法在图像质量与偏好对齐上显著优于强基线,验证了表征提取与对齐的有效性。

原文摘要 · Abstract (English)

Preference-conditioned image generation seeks to adapt generative models to individual users, producing outputs that reflect personal aesthetic choices beyond the given textual prompt. Despite recent progress, existing approaches either fail to capture nuanced user preferences or lack effective mechanisms to encode personalized visual signals. In this work, we propose a multimodal framework that leverages multimodal large language models (MLLMs) to extract rich user representations and inject them into diffusion-based image generation. We train the MLLM with a preference-oriented visual question answering task to capture fine-grained semantic cues. To isolate preference-relevant features, we introduce two complementary probing tasks: inter-user discrimination to distinguish between different users, and intra-user discrimination to separate liked from disliked content. To ensure compatibility with diffusion text encoders, we design a maximum mean discrepancy-based alignment loss that bridges the modality gap while preserving multimodal structure. The resulting embeddings are used to condition the generator, enabling faithful adherence to both prompts and user preferences. Extensive experiments demonstrate that our method substantially outperforms strong baselines in both image quality and preference alignment, highlighting the effectiveness of representation extraction and alignment for personalized generation.

图像生成多模态个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。