arXiv:2510.05283cs.AIcs.CL2025-10

用混合奖励机制提升多模态模型对齐效果

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

  • 融合学习型与规则型奖励,兼顾灵活性与准确性
  • 数学推理任务平均提升16%,整体任务提升9.5%
  • 适合需要高可靠对齐的多模态应用开发者

多模态大语言模型(MLLM)对齐人类偏好通常依赖单一信号、基于模型的奖励方法。这类单体奖励在特定任务中缺乏置信度校准,难以捕捉人类偏好的多维度特征,且需大量数据标注和奖励模型训练。本文提出一种混合奖励建模范式,整合两类互补机制:(i) 基于模型的奖励,通过合成与人工反馈学习预测标量或向量得分;(ii) 基于规则的奖励,利用领域特定启发式提供明确正确性信号及置信度。除准确率外,进一步引入多方面奖励以强化指令遵循,并设计通用长度惩罚奖励以稳定训练并提升性能。该框架通过强化学习策略优化,实现对MLLM的有效对齐。实验表明,在多个多模态基准上,混合与多方面奖励建模均带来一致改进。3B规模模型在通用与数学推理任务上平均提升约9.5%;聚焦数学基准时,平均提升达约16%,凸显其在数学推理与问题求解中的有效性。

原文摘要 · Abstract (English)

Aligning multimodal large language models (MLLMs) with human preferences often relies on single-signal, model-based reward methods. Such monolithic rewards often lack confidence calibration across domain-specific tasks, fail to capture diverse aspects of human preferences, and require extensive data annotation and reward model training. In this work, we propose a hybrid reward modeling framework that integrates complementary reward paradigms: (i) model-based rewards, where a learned reward model predicts scalar or vector scores from synthetic and human feedback, and (ii) rule-based rewards, where domain-specific heuristics provide explicit correctness signals with confidence. Beyond accuracy, we further incorporate multi-aspect rewards to enforce instruction adherence and introduce a generalized length-penalty reward to stabilize training and improve performance. The proposed framework provides a flexible and effective approach to aligning MLLMs through reinforcement learning policy optimization. Our experiments show consistent improvements across different multimodal benchmarks when applying hybrid and multi-aspect reward modeling. Our best performing model in the 3B family achieves an overall average improvement of ~9.5% across general and math reasoning tasks. Focusing specifically on mathematical benchmarks, the model achieves a significant average improvement of ~16%, highlighting its effectiveness in mathematical reasoning and problem solving.

多模态对齐奖励建模数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。