arXiv:2602.00846cs.CL2026-02被引 7

构建多模态奖励模型,自动合成高质量偏好数据提升对齐效果

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

  • 用规则引导自动生成跨模态偏好数据,替代昂贵人工标注
  • 在视频和音频任务上分别达到80.2%和66.8%准确率,整体提升17%
  • 支持多模态推理且可迁移至纯文本任务,适合对齐研究者使用

多模态大语言模型因现有奖励模型多以视觉为中心、依赖高成本人工标注、输出不透明的标量分数而难以实现良好对齐。本文提出Omni-RRM,一种基于规则的多模态奖励模型,可生成文本、图像、视频和音频的多维奖励信号。为克服人工评估成本高且不一致的问题,提出自动规则引导偏好合成方法(Omni-Preference),通过教师模型将原始偏好转化为明确解释,确保合成监督的高保真与可解释性。Omni-RRM采用渐进式SFT+GRPO训练策略,专门优化低置信度难样本的判别能力。在视频(ShareGPT-Video上80.2%)、音频(Audio-HH-RLHF上66.8%,TA2T上65.0%)等五项基准上综合准确率达70.4%,相较基线提升17.0%。该模型有效指导Best-of-N选择,并在纯文本对齐任务中表现出强泛化能力。全部资源已开源。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar scores that fail to capture nuanced reasoning, leading to brittle alignment. We present Omni-RRM, an \textbf{Omni}-modal \textbf{R}ubric-grounded \textbf{R}eward \textbf{M}odel that generates multi-dimensional reward signals across text, image, video, and audio. To overcome the high cost and inherent inconsistency of human-centric evaluation in multi-dimensional reasoning, we introduce \textbf{Omni-Preference}, a high-quality dataset constructed via automatic rubric-grounded preference synthesis. In this pipeline, teacher models reconcile raw preferences into explicit justifications, ensuring that the synthesized supervision is both high-fidelity and interpretable. Omni-RRM is trained using a progressive SFT + GRPO regimen, specifically optimized to sharpen reward discrimination on low-margin, hard preference pairs. It achieves state-of-the-art accuracy on video (80.2\% on ShareGPT-Video) and audio benchmarks (66.8\% on Audio-HH-RLHF and 65.0\% on TA2T), yielding a five-benchmark Overall accuracy of 70.4\% and a +17.0\% relative gain over its backbone. Furthermore, Omni-RRM effectively guides Best-of-$N$ selection and exhibits robust transfer to text-only alignment. All resources, including the dataset, training and inference code, and model checkpoints are available at https://tmfk418.github.io/Omni-RRM.

多模态奖励建模自动化标注对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。