构建支持多模态自由偏好评分的通用奖励模型
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- 提出跨文本、图像、视频、音频等9类任务的多模态评分基准
- 构建含24.8万条通用偏好对的多模态数据集
- 支持自由形式偏好,适合需要个性化对齐的AI系统
奖励模型(RMs)在对齐人工智能行为与人类偏好方面起关键作用,但面临两大挑战:(1) 模态不平衡,现有模型主要聚焦于文本和图像,对视频、音频等模态支持有限;(2) 偏好僵化,基于固定二元偏好对训练难以捕捉个性化偏好的复杂性与多样性。为应对上述问题,我们提出Omni-Reward,迈向支持自由形式偏好的通用多模态奖励建模:(1) 评估:引入Omni-RewardBench,首个包含自由形式偏好的多模态奖励模型评测基准,涵盖五种模态(文本、图像、视频、音频、3D)下的九项任务;(2) 数据:构建Omni-RewardData,一个包含24.8万条通用偏好对和6.9万条指令微调对的多模态偏好数据集;(3) 模型:提出Omni-RewardModel,包含判别式与生成式奖励模型,在Omni-RewardBench及其他主流奖励建模基准上均表现优异。
原文摘要 · Abstract (English)
Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and (2) Preference Rigidity, where training on fixed binary preference pairs fails to capture the complexity and diversity of personalized preferences. To address the above challenges, we propose Omni-Reward, a step toward generalist omni-modal reward modeling with support for free-form preferences, consisting of: (1) Evaluation: We introduce Omni-RewardBench, the first omni-modal RM benchmark with free-form preferences, covering nine tasks across five modalities including text, image, video, audio, and 3D; (2) Data: We construct Omni-RewardData, a multimodal preference dataset comprising 248K general preference pairs and 69K instruction-tuning pairs for training generalist omni-modal RMs; (3) Model: We propose Omni-RewardModel, which includes both discriminative and generative RMs, and achieves strong performance on Omni-RewardBench as well as other widely used reward modeling benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。