评测多模态大模型在跨模态知识推理中的一致性问题
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
- 设计四类评测任务,构建新数据集评估多模态推理一致性
- 发现当前多模态大模型在跨模态推理中存在显著一致性退化
- 揭示影响推理一致性的关键因素,为模型优化提供方向
近年来,多模态大语言模型(MLLMs)在文本与视觉理解方面取得显著进展。然而,当前MLLMs在多模态知识推理过程中仍难以有效融合跨模态知识,导致推理结果不一致。为系统探究该问题,我们提出了四项评估任务,并构建了一个新数据集。基于该数据集,我们开展了一系列实验,分析并比较了不同MLLM在多模态知识推理中的一致性退化程度。根据实验结果,我们识别出导致一致性下降的关键因素。本研究为多模态知识推理的挑战提供了新见解,并为未来提升MLLM性能提供了重要参考。
原文摘要 · Abstract (English)
In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating knowledge across these modalities during multimodal knowledge reasoning, leading to inconsistencies in reasoning outcomes. To systematically explore this issue, we propose four evaluation tasks and construct a new dataset. We conduct a series of experiments on this dataset to analyze and compare the extent of consistency degradation in multimodal knowledge reasoning within MLLMs. Based on the experimental results, we identify factors contributing to the observed degradation in consistency. Our research provides new insights into the challenges of multimodal knowledge reasoning and offers valuable guidance for future efforts aimed at improving MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。