arXiv:2410.15334cs.CV2024-10IJCAI被引 29

提升多模态大模型可信度,让回答更忠实于图像内容。

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

  • 构建仅在关键区域不同的图像偏好数据集,精准定位模态偏差。
  • 引入图像奖励损失函数,使模型回答更贴合输入图像。
  • 迭代式对齐策略稳定训练,7B模型达13B以上水平。

多模态大语言模型(MLLMs)在各类任务中表现卓越,但视觉与文本编码器的独立训练常导致模态错位,引发幻觉问题——模型生成输入图像中不存在的内容,严重影响其在真实场景中的可信度。尽管已有优化文本偏好的方法,我们发现模型在图像严重失真时仍倾向于输出偏好答案,且注意力集中于上下文而非关键对象。为此,我们提出模态公平偏好优化(MFPO),包含三部分:构建仅在关键区域差异的多模态偏好数据集;设计图像奖励损失函数,强化回答与图像的一致性;采用由易到难的迭代对齐策略,稳定联合训练。在三个可信度基准上的实验表明,MFPO显著提升模型可信度,使7B模型达到甚至超越13B、34B等更大模型的水平。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models.

多模态可信度偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。