arXiv:2602.02380cs.CV2026-02被引 15

让视觉生成更懂人心,个性化评分模型提升内容匹配度

Unified Personalized Reward Model for Vision Generation

  • 结合语义理解与动态评估框架,按上下文灵活生成评分标准
  • 在图像和视频生成中显著提升对主观偏好的一致性表现
  • 适合需要高精度人感对齐的生成任务,如艺术创作与广告设计

多模态奖励模型(RMs)的进展推动了视觉生成的发展。现有方法多采用布拉德利-特瑞样式偏好建模或使用生成式视觉语言模型(VLM)作为评判者,并通过强化学习优化生成模型。然而,当前RMs存在固有局限:通常采用‘一刀切’范式,假设统一的偏好分布或依赖固定评价标准,难以捕捉特定内容的视觉线索,导致与主观、上下文相关的用户偏好系统性偏离。为此,受人类评估启发,我们提出UnifiedReward-Flex——一种统一的个性化视觉生成奖励模型,将奖励建模与灵活自适应推理相结合。给定提示和生成内容后,模型先解析语义意图并基于视觉证据进行定位,再动态构建分层评估体系,在预定义与自生成的高层维度下实例化细粒度评判标准。训练流程分两阶段:(1) 从先进闭源VLM中蒸馏结构化高质量推理轨迹,用于监督微调(SFT),赋予模型灵活自适应推理能力;(2) 在精心筛选的偏好对上执行直接偏好优化(DPO),进一步增强推理准确性和判别一致性。为验证有效性,我们将UnifiedReward-Flex集成至GRPO框架用于图像与视频合成,实验结果证明其显著优于基线。

原文摘要 · Abstract (English)

Recent advancements in multimodal reward models (RMs) have significantly propelled the development of visual generation. Existing frameworks typically adopt Bradley-Terry-style preference modeling or leverage generative VLMs as judges, and subsequently optimize visual generation models via reinforcement learning. However, current RMs suffer from inherent limitations: they often follow a one-size-fits-all paradigm that assumes a monolithic preference distribution or relies on fixed evaluation rubrics. As a result, they are insensitive to content-specific visual cues, leading to systematic misalignment with subjective and context-dependent human preferences. To this end, inspired by human assessment, we propose UnifiedReward-Flex, a unified personalized reward model for vision generation that couples reward modeling with flexible and context-adaptive reasoning. Specifically, given a prompt and the generated visual content, it first interprets the semantic intent and grounds on visual evidence, then dynamically constructs a hierarchical assessment by instantiating fine-grained criteria under both predefined and self-generated high-level dimensions. Our training pipeline follows a two-stage process: (1) we first distill structured, high-quality reasoning traces from advanced closed-source VLMs to bootstrap SFT, equipping the model with flexible and context-adaptive reasoning behaviors; (2) we then perform direct preference optimization (DPO) on carefully curated preference pairs to further strengthen reasoning fidelity and discriminative alignment. To validate the effectiveness, we integrate UnifiedReward-Flex into the GRPO framework for image and video synthesis, and extensive results demonstrate its superiority.

视觉生成个性化评分奖励模型自适应推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。