arXiv:2503.03122cs.CLcs.AI2025-03

解决多模态奖励模型依赖文本捷径的问题,提升泛化能力。

The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models

  • 通过动态重加权样本,减少对单模态文本捷径的依赖。
  • 在分布外数据上显著提升泛化性能与下游任务表现。
  • 适合需要可靠多模态对齐的LLM应用开发人员。

多模态奖励模型(MM-RMs)对齐大语言模型(LLMs)与人类偏好至关重要,尤其当LLMs越来越多地处理多模态数据时。然而,我们发现现有数据集训练的MM-RMs往往难以泛化到分布外数据,原因是其过度依赖单模态伪相关性,主要是训练分布中的纯文本捷径,从而无法真正利用多模态奖励函数。为此,我们提出一种感知捷径的MM-RM学习算法,通过动态重加权训练样本,使分布更倾向于真正的多模态理解,降低对单模态伪相关性的依赖。实验表明,该方法在泛化性、下游任务性能和可扩展性方面均有显著提升,建立了一个更鲁棒的多模态奖励建模框架。

原文摘要 · Abstract (English)

Multimodal Reward Models (MM-RMs) are crucial for aligning Large Language Models (LLMs) with human preferences, particularly as LLMs increasingly interact with multimodal data. However, we find that MM-RMs trained on existing datasets often struggle to generalize to out-of-distribution data due to their reliance on unimodal spurious correlations, primarily text-only shortcuts within the training distribution, which prevents them from leveraging true multimodal reward functions. To address this, we introduce a Shortcut-aware MM-RM learning algorithm that mitigates this issue by dynamically reweighting training samples, shifting the distribution toward better multimodal understanding, and reducing dependence on unimodal spurious correlations. Our experiments demonstrate significant improvements in generalization, downstream task performance, and scalability, establishing a more robust framework for multimodal reward modeling.

多模态奖励模型泛化性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。