调整输入顺序可显著提升多模态模型性能,最高增益达17.8%
Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
- 通过将关键内容置于输入序列首尾位置,利用模型对位置的敏感性
- 视频描述匹配任务提升14.7%,视觉问答任务提升17.8%
- 提出新评估指标PIA,缓解模型对输入顺序的依赖偏差
多模态大语言模型(MLLMs)利用包含文本、图像或视频的多模态上下文解决各类任务。我们发现,改变多模态输入顺序会导致模型性能在先进表现与随机猜测间剧烈波动,该现象存在于纯文本、纯图像及图文混合场景中。进一步研究表明,主流MLLM对上下文特定位置(尤其是开头和结尾)具有特殊关注。基于此,我们将关键视频帧及重要图文内容置于上下文特殊位置进行推理,使视频-描述匹配任务平均提升14.7%,视觉问答任务提升17.8%。此外,我们提出新的评估指标——位置不变准确率(PIA),以应对MLLM评估中的顺序偏差问题。研究结果深化了对多模态上下文学习(MMICL)的理解,并提供了无需增加计算成本的实用性能优化策略。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) utilize multimodal contexts consisting of text, images, or videos to solve various multimodal tasks. However, we find that changing the order of multimodal input can cause the model's performance to fluctuate between advanced performance and random guessing. This phenomenon exists in both single-modality (text-only or image-only) and mixed-modality (image-text-pair) contexts. Furthermore, we demonstrate that popular MLLMs pay special attention to certain multimodal context positions, particularly the beginning and end. Leveraging this special attention, we place key video frames and important image/text content in special positions within the context and submit them to the MLLM for inference. This method results in average performance gains of 14.7% for video-caption matching and 17.8% for visual question answering tasks. Additionally, we propose a new metric, Position-Invariant Accuracy (PIA), to address order bias in MLLM evaluation. Our research findings contribute to a better understanding of Multi-Modal In-Context Learning (MMICL) and provide practical strategies for enhancing MLLM performance without increasing computational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。