arXiv:2604.19544cs.AI2026-04

提出新方法提升多模态奖励模型训练数据质量,性能达新高。

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

论文配图:DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling
图 1 · 摘自论文原文
  • 去偏偏好构建+文本到图像数据重构+迭代训练,系统优化数据质量。
  • 在三个主流评测集上达到当前最优表现,显著优于现有方法。
  • 适合研究多模态对齐、奖励建模或数据清洗的从业者使用。

多模态奖励模型(MRMs)在对齐多模态大语言模型(MLLMs)与人类偏好方面起关键作用。高质量的多模态偏好数据是训练有效MRM的基础。然而,现有偏好数据集存在三大挑战:偏好强度粒度不足、文本风格偏差以及不可靠的偏好信号。此外,开源多模态偏好数据普遍噪声严重,但缺乏高效可扩展的数据清洗方法。为此,我们提出DT2IT-MRM,融合去偏偏好构建流程、新颖的文本到图像(T2I)偏好数据重构方式,以及用于多模态奖励建模的迭代训练框架,以优化现有数据集质量。实验表明,DT2IT-MRM在三个主要基准测试——VL-RewardBench、Multimodal RewardBench和MM-RLHF-RewardBench上均达到新的最佳性能。

原文摘要 · Abstract (English)

Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM requires high-quality multimodal preference data. However, existing preference datasets face three key challenges: lack of granularity in preference strength, textual style bias, and unreliable preference signals. Besides, existing open-source multimodal preference datasets suffer from substantial noise, yet there is a lack of effective and scalable curation methods to enhance their quality. To address these limitations, we propose \textbf{DT2IT-MRM}, which integrates a \textbf{D}ebiased preference construction pipeline, a novel reformulation of text-to-image (\textbf{T2I}) preference data, and an \textbf{I}terative \textbf{T}raining framework that curates existing multimodal preference datasets for \textbf{M}ultimodal \textbf{R}eward \textbf{M}odeling. Our experimental results show that DT2IT-MRM achieves new \textbf{state-of-the-art} overall performance on three major benchmarks: VL-RewardBench, Multimodal RewardBench, and MM-RLHF-RewardBench.

多模态奖励模型数据清洗对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。