让生成内容更懂用户偏好,图文一致且个性化。
Discrete Preference Learning for Personalized Multimodal Generation

- 用图神经网络学习用户对图文的独立偏好,转为离散编码
- 在图文生成中注入偏好编码,提升个性化与一致性
- 适合需要精准定制多模态内容的研究与应用
生成模型可生成符合用户偏好的文本和图像。现有个性化生成模型存在两大缺陷:缺乏专门的偏好建模范式,且仅生成单模态内容,无法反映真实世界中多模态交互行为。为此,我们提出个性化多模态生成方法,通过多模态交互构建专用偏好模型,捕捉模态特异性偏好,并将其输入下游生成器以生成个性化多模态内容。该任务面临两大挑战:(1)专用建模得到的连续偏好与生成器固有的离散标记输入之间存在鸿沟;(2)生成图文间可能存在不一致。为此,我们提出两阶段框架DPPMG。第一阶段,引入模态特异性图神经网络(专用偏好模型)学习用户偏好,并将偏好量化为离散偏好标记。第二阶段,将这些离散偏好标记注入下游文本与图像生成器。为进一步增强跨模态一致性并保留个性化,设计跨模态一致且个性化的奖励信号,微调标记相关参数。在两个真实数据集上的大量实验表明,该模型能有效生成个性化且一致的多模态内容。
原文摘要 · Abstract (English)
The emergence of generative models enables the creation of texts and images tailored to users' preferences. Existing personalized generative models have two critical limitations: lacking a dedicated paradigm for accurate preference modeling, and generating unimodal content despite real-world multimodal-driven user interactions. Therefore, we propose personalized multimodal generation, which captures modal-specific preferences via a dedicated preference model from multimodal interactions, and then feeds them into downstream generators for personalized multimodal content. However, this task presents two challenges: (1) Gap between continuous preferences from dedicated modeling and discrete token inputs intrinsic to generator architectures; (2) Potential inconsistency between generated images and texts. To tackle these, we present a two-stage framework called Discrete Preference learning for Personalized Multimodal Generation (DPPMG). In the first stage, to accurately learn discrete modal-specific preferences, we introduce a modal-specific graph neural network (a dedicated preference model) to learn users' modal-specific preferences, which preferences are then quantized into discrete preference tokens. In the second stage, the discrete modal-specific preference tokens are injected into downstream text and image generators. To further enhance cross-modal consistency while preserving personalization, we design a cross-modal consistent and personalized reward to fine-tune token-associated parameters. Extensive experiments on two real-world datasets demonstrate the effectiveness of our model in generating personalized and consistent multimodal content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。