提出新方法提升多模态奖励模型训练数据质量,性能达新高。
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

- 去偏偏好构建+文本到图像数据重构+迭代训练,系统优化数据质量。
- 在三个主流评测集上达到当前最优表现,显著优于现有方法。
- 适合研究多模态对齐、奖励建模或数据清洗的从业者使用。
多模态奖励模型(MRMs)在对齐多模态大语言模型(MLLMs)与人类偏好方面起关键作用。高质量的多模态偏好数据是训练有效MRM的基础。然而,现有偏好数据集存在三大挑战:偏好强度粒度不足、文本风格偏差以及不可靠的偏好信号。此外,开源多模态偏好数据普遍噪声严重,但缺乏高效可扩展的数据清洗方法。为此,我们提出DT2IT-MRM,融合去偏偏好构建流程、新颖的文本到图像(T2I)偏好数据重构方式,以及用于多模态奖励建模的迭代训练框架,以优化现有数据集质量。实验表明,DT2IT-MRM在三个主要基准测试——VL-RewardBench、Multimodal RewardBench和MM-RLHF-RewardBench上均达到新的最佳性能。
原文摘要 · Abstract (English)
Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM requires high-quality multimodal preference data. However, existing preference datasets face three key challenges: lack of granularity in preference strength, textual style bias, and unreliable preference signals. Besides, existing open-source multimodal preference datasets suffer from substantial noise, yet there is a lack of effective and scalable curation methods to enhance their quality. To address these limitations, we propose \textbf{DT2IT-MRM}, which integrates a \textbf{D}ebiased preference construction pipeline, a novel reformulation of text-to-image (\textbf{T2I}) preference data, and an \textbf{I}terative \textbf{T}raining framework that curates existing multimodal preference datasets for \textbf{M}ultimodal \textbf{R}eward \textbf{M}odeling. Our experimental results show that DT2IT-MRM achieves new \textbf{state-of-the-art} overall performance on three major benchmarks: VL-RewardBench, Multimodal RewardBench, and MM-RLHF-RewardBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。