将隐式偏好转化为可解释的多维度评分标准,提升生成模型对齐效果。
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria

- 通过自动提取视觉语言模型内部偏好,生成可验证的评分规则
- 在图文生成与图像编辑任务中优于传统成对打分与模型评判
- 适合需要可解释性、低样本训练的生成模型对齐场景
对齐多模态生成模型与人类偏好,需依赖能反映判断多维性与组合性的奖励信号。现有强化学习人类反馈(RLHF)方法将其简化为标量或成对标签,导致细粒度偏好丢失,并易受奖励劫持。尽管近期提出的‘评分标准即奖励’(RaR)方法试图恢复这一结构,但如何生成既可靠又高效、可扩展的评分标准仍是难题。本文提出自动评分标准作为奖励(ARR),将奖励建模从隐式权重优化转变为显式的、基于准则的分解。在任何成对比较前,ARR将视觉语言模型(VLM)内化的偏好知识外化为特定提示的评分标准,将整体意图转化为独立可验证的质量维度。这一转换使评估偏差(如位置偏差)显著降低,支持零样本部署与少样本条件化。为进一步用于生成训练,我们提出评分标准策略优化(RPO),将ARR的多维评估凝练为鲁棒的二元奖励,以评分条件化的偏好决策替代模糊的标量回归,稳定策略梯度。在文本到图像生成和图像编辑基准上,ARR-RPO优于成对奖励模型与VLM裁判,表明显式外化隐式偏好知识为结构化评分标准,可实现更可靠、数据高效的多模态对齐,揭示瓶颈在于缺乏因子化接口,而非知识不足。
原文摘要 · Abstract (English)
Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human judgment. Prevailing RLHF approaches reduce this structure to scalar or pairwise labels, collapsing nuanced preferences into opaque parametric proxies and exposing vulnerabilities to reward hacking. While recent Rubrics-as-Reward (RaR) methods attempt to recover this structure through explicit criteria, generating rubrics that are simultaneously reliable, scalable, and data-efficient remains an open problem. We introduce Auto-Rubric as Reward (ARR), a framework that reframes reward modeling from implicit weight optimization to explicit, criteria-based decomposition. Before any pairwise comparison, ARR externalizes a VLM's internalized preference knowledge as prompt-specific rubrics, translating holistic intent into independently verifiable quality dimensions. This conversion of implicit preference structure into inspectable, interpretable constraints substantially suppresses evaluation biases including positional bias, enabling both zero-shot deployment and few-shot conditioning on minimal supervision. To extend these gains into generative training, we propose Rubric Policy Optimization (RPO), which distills ARR's structured multi-dimensional evaluation into a robust binary reward, replacing opaque scalar regression with rubric-conditioned preference decisions that stabilize policy gradients. On text-to-image generation and image editing benchmarks, ARR-RPO outperforms pairwise reward models and VLM judges, demonstrating that explicitly externalizing implicit preference knowledge into structured rubrics achieves more reliable, data-efficient multimodal alignment, revealing that the bottleneck is the absence of a factorized interface, not a deficit of knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。