发现视觉语言模型在跨模态任务中表现不一致,揭示其融合缺陷。
Cross-Modal Consistency in Multimodal Large Language Models
- 提出跨模态一致性概念,量化视觉与语言模态表现差异
- 在自建多模态数据集上测试,发现GPT-4V存在显著模态不一致
- 为模型设计优化提供依据,适合关注多模态融合的研究者
近期多模态方法的发展开启了处理文本、音频和视觉内容等多样化数据的新纪元。如GPT-4V这类将计算机视觉与先进语言处理结合的模型,在需同时理解文本与视觉信息的复杂任务中表现出卓越能力。以往研究虽在目标检测、图像描述等场景对视觉大语言模型(VLLMs)进行了细致评估,但大多仅孤立分析各模态性能,忽视了模态间的复杂交互。尤其未回答:同一任务实例在不同模态下模型是否保持相同准确率。本文首次深入探究这些模态间的互动与对比,引入‘跨模态一致性’新概念,并构建基于该概念的定量评估框架。基于我们自建的平行视觉-语言数据集,实验结果揭示,尽管被视作统一多模态模型,GPT-4V在视觉与语言模态间仍存在显著不一致。本研究为模型的合理使用提供了洞见,并指明了改进方向。
原文摘要 · Abstract (English)
Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision with advanced language processing, exhibit extraordinary proficiency in handling intricate tasks that require a simultaneous understanding of both textual and visual information. Prior research efforts have meticulously evaluated the efficacy of these Vision Large Language Models (VLLMs) in various domains, including object detection, image captioning, and other related fields. However, existing analyses have often suffered from limitations, primarily centering on the isolated evaluation of each modality's performance while neglecting to explore their intricate cross-modal interactions. Specifically, the question of whether these models achieve the same level of accuracy when confronted with identical task instances across different modalities remains unanswered. In this study, we take the initiative to delve into the interaction and comparison among these modalities of interest by introducing a novel concept termed cross-modal consistency. Furthermore, we propose a quantitative evaluation framework founded on this concept. Our experimental findings, drawn from a curated collection of parallel vision-language datasets developed by us, unveil a pronounced inconsistency between the vision and language modalities within GPT-4V, despite its portrayal as a unified multimodal model. Our research yields insights into the appropriate utilization of such models and hints at potential avenues for enhancing their design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。