arXiv:2411.01307cs.CL2024-11被引 2

探究多模态大模型能否类比推理,验证其理解与预测能力。

Can Multimodal Large Language Model Think Analogically?

  • 设计统一提示模板,挖掘模型对类比问题的理解力。
  • 在主流数据集上表现优于现有方法,支持类比推理能力存在。
  • 适合关注多模态推理、模型认知能力的研究者参考。

类比推理是人类感知与创造力的基础,尤其在多模态情境中尤为重要。近年来,多模态大语言模型(MLLM)因其涌现能力引发广泛关注。本文深入探讨了MLLM在多模态类比推理方面的能力,分为两个层面: 1. MLLM作为解释者:研究其是否能深入理解多模态类比推理问题。提出统一的提示模板及利用模型理解能力增强现有模型的方法。 2. MLLM作为预测者:探究其能否直接解决多模态类比推理任务。实验表明,所提方法在多个主流数据集上优于现有技术,为MLLM具备类比推理能力提供了初步证据。

原文摘要 · Abstract (English)

Analogical reasoning, particularly in multimodal contexts, is the foundation of human perception and creativity. Multimodal Large Language Model (MLLM) has recently sparked considerable discussion due to its emergent capabilities. In this paper, we delve into the multimodal analogical reasoning capability of MLLM. Specifically, we explore two facets: \textit{MLLM as an explainer} and \textit{MLLM as a predictor}. In \textit{MLLM as an explainer}, we primarily focus on whether MLLM can deeply comprehend multimodal analogical reasoning problems. We propose a unified prompt template and a method for harnessing the comprehension capabilities of MLLM to augment existing models. In \textit{MLLM as a predictor}, we aim to determine whether MLLM can directly solve multimodal analogical reasoning problems. The experiments show that our approach outperforms existing methods on popular datasets, providing preliminary evidence for the analogical reasoning capability of MLLM.

多模态类比推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。