用响应熵筛选数据并分阶段训练,提升多模态推理奖励模型效率
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
- 基于响应熵识别难样本和噪声数据,自动筛选优质训练样本
- 在三个基准上性能超越现有模型,显著提升准确率
- 适合需要高效训练多模态推理系统的研究者使用
多模态奖励模型对对齐多模态大语言模型与人类偏好至关重要。近期工作已将推理能力融入此类模型,取得良好效果。然而,训练过程面临两大挑战:(1) 偏好数据集固有的噪声会降低模型性能;(2) 传统训练方法忽略样本难度差异,效率低下。本文发现响应熵与准确率存在强相关性,表明熵可作为无监督的标注噪声与样本难度代理指标。基于此,提出熵引导训练(EGT)方法,包含两项策略:(1) 基于熵的数据清洗,减少不可靠样本影响;(2) 按熵值渐进引入更复杂样本。在三个基准上的大量实验表明,经EGT训练的模型始终优于当前最优多模态奖励模型。
原文摘要 · Abstract (English)
Multimodal reward models are crucial for aligning multimodal large language models with human preferences. Recent works have incorporated reasoning capabilities into these models, achieving promising results. However, training these models suffers from two critical challenges: (1) the inherent noise in preference datasets, which degrades model performance, and (2) the inefficiency of conventional training methods, which ignore the differences in sample difficulty. In this paper, we identify a strong correlation between response entropy and accuracy, indicating that entropy can serve as a reliable and unsupervised proxy for annotation noise and sample difficulty. Based on this insight, we propose a novel Entropy-Guided Training (EGT) approach for multimodal reasoning reward models, which combines two strategies: (1) entropy-guided data curation to mitigate the impact of unreliable samples, and (2) an entropy-guided training strategy that progressively introduces more complex examples. Extensive experiments across three benchmarks show that the EGT-trained model consistently outperforms state-of-the-art multimodal reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。