arXiv:2409.19339cs.CLcs.AI2024-09EMNLP被引 11

提升多模态大模型的视觉问题分解能力,让模型更聪明地拆解复杂问题。

Visual Question Decomposition on Multimodal Large Language Models

  • 构建评估框架与数据集,系统评测多模态模型的问题分解质量。
  • 提出DecoVQA+数据集和高效微调方法,显著提升子问题质量。
  • 模型在视觉问答任务中实现更高准确率,适合需要精准推理的研究者。

问题分解已成为提升大型语言模型回答复杂问题能力的有效策略。然而,现有方法主要针对单模态语言模型,多模态大模型(MLLMs)的问题分解能力尚未被充分探索。为此,本文系统研究了多模态大模型的视觉问题分解能力。我们构建了一个包含数据集和多项评估标准的评估框架,揭示了现有MLLMs难以生成高质量子问题。为解决此问题,我们提出了一个专门的微调数据集DecoVQA+,并设计了一种高效的微调流程,旨在使模型能够进行适当的有选择性分解。该流程结合所提数据集与选择性分解训练目标。经过微调的MLLM在子问题质量和选择性分解策略上均有显著提升,并在VQA基准数据集上实现了更高的准确性。

原文摘要 · Abstract (English)

Question decomposition has emerged as an effective strategy for prompting Large Language Models (LLMs) to answer complex questions. However, while existing methods primarily focus on unimodal language models, the question decomposition capability of Multimodal Large Language Models (MLLMs) has yet to be explored. To this end, this paper explores visual question decomposition on MLLMs. Specifically, we introduce a systematic evaluation framework including a dataset and several evaluation criteria to assess the quality of the decomposed sub-questions, revealing that existing MLLMs struggle to produce high-quality sub-questions. To address this limitation, we propose a specific finetuning dataset, DecoVQA+, for enhancing the model's question decomposition capability. Aiming at enabling models to perform appropriate selective decomposition, we propose an efficient finetuning pipeline. The finetuning pipeline consists of our proposed dataset and a training objective for selective decomposition. Finetuned MLLMs demonstrate significant improvements in the quality of sub-questions and the policy of selective question decomposition. Additionally, the models also achieve higher accuracy with selective decomposition on VQA benchmark datasets.

多模态问题分解大模型视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。