为文生图设计个性化评分模型,让生成图像更贴合个人喜好。
Personalized Reward Modeling for Text-to-Image Generation
- 基于用户数据动态生成评价维度,用思维链推理评估图像
- 仅需少量参考数据即可构建完整用户画像,无需单独训练
- 能给出定制化反馈,帮助优化提示词,适合个性化需求者
近期文本到图像(T2I)模型能从文本提示生成语义一致的图像,但如何评估生成结果与个体用户偏好的对齐程度仍是难题。传统方法使用通用奖励函数或相似性度量,难以捕捉个人视觉偏好的多样性与复杂性。本文提出PIGReward,一种个性化奖励模型,通过思维链(CoT)推理动态生成用户相关的评估维度,并在有限参考数据下采用自举策略构建丰富用户上下文,实现无需用户专属训练的个性化评估。该模型不仅可精准评分,还能提供个性化反馈以驱动用户特定提示词优化,提升生成内容与个人意图的一致性。我们进一步构建了PIGBench,一个针对用户偏好的基准测试,涵盖共享提示下的多样视觉表达。大量实验表明,PIGReward在准确性和可解释性上均优于现有方法,为可扩展、基于推理的个性化T2I评估与优化奠定了基础。
原文摘要 · Abstract (English)
Recent text-to-image (T2I) models generate semantically coherent images from textual prompts, yet evaluating how well they align with individual user preferences remains an open challenge. Conventional evaluation methods, general reward functions or similarity-based metrics, fail to capture the diversity and complexity of personal visual tastes. In this work, we present PIGReward, a personalized reward model that dynamically generates user-conditioned evaluation dimensions and assesses images through CoT reasoning. To address the scarcity of user data, PIGReward adopt a self-bootstrapping strategy that reasons over limited reference data to construct rich user contexts, enabling personalization without user-specific training. Beyond evaluation, PIGReward provides personalized feedback that drives user-specific prompt optimization, improving alignment between generated images and individual intent. We further introduce PIGBench, a per-user preference benchmark capturing diverse visual interpretations of shared prompts. Extensive experiments demonstrate that PIGReward surpasses existing methods in both accuracy and interpretability, establishing a scalable and reasoning-based foundation for personalized T2I evaluation and optimization. Taken together, our findings highlight PIGReward as a robust steptoward individually aligned T2I generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。