让视觉生成模型关注重要区域,提升细节质量与人类偏好对齐
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- 将图像/视频的奖励拆解为像素级优势图,精准定位关键区域
- 在图像和视频任务中均超越传统方法,提升对齐度与泛化能力
- 无需改动现有训练流程,轻量高效适配各类生成模型
强化学习已成为后训练视觉生成模型的强大工具,其中群体相对策略优化(GRPO)被广泛用于对齐生成结果与人类偏好。然而,现有GRPO流程对每个样本仅使用单一标量奖励,将图像或视频视为整体,忽略了视觉内容丰富的空间与时间结构。这种粗粒度监督限制了局部伪影的修正和细微感知线索的建模。本文提出视觉偏好策略优化(ViPO),一种提升标量反馈为结构化像素级优势的GRPO变体。ViPO引入感知结构模块,利用预训练视觉骨干网络构建时空感知的优势图,将优化压力集中于感知重要区域,同时保持标准GRPO的稳定性。在图像与视频基准测试中,ViPO持续优于原始GRPO,提升了域内人类偏好对齐效果,并增强了域外评估的泛化能力。该方法架构无关、轻量且完全兼容现有GRPO训练流程,为视觉生成提供了更丰富、更精确的学习信号。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each image or video as a holistic entity and ignoring the rich spatial and temporal structure of visual content. This coarse supervision hinders the correction of localized artifacts and the modeling of fine-grained perceptual cues. We introduce Visual Preference Policy Optimization (ViPO), a GRPO variant that lifts scalar feedback into structured, pixel-level advantages. ViPO employs a Perceptual Structuring Module that uses pretrained vision backbones to construct spatially and temporally aware advantage maps, redistributing optimization pressure toward perceptually important regions while preserving the stability of standard GRPO. Across both image and video benchmarks, ViPO consistently outperforms vanilla GRPO, improving in-domain alignment with human-preference rewards and enhancing generalization on out-of-domain evaluations. The method is architecture-agnostic, lightweight, and fully compatible with existing GRPO training pipelines, providing a more expressive and informative learning signal for visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。