让视觉语言模型像人一样迭代思考,提升复杂指令理解能力
CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs
- 引入上下文感知的迭代推理与自我修正机制
- 在多模态任务上达到91.5%准确率,超越GPT-4V等主流模型
- 适合需要逻辑推理与动态反馈的复杂多模态应用
大型语言模型(LLMs)和大型视觉语言模型(LVLMs)的发展显著提升了对文本和视觉信息的处理与生成能力。然而,这些模型在处理需要逻辑推理、动态反馈整合与迭代自我修正的复杂多步多模态指令时仍表现不佳。为此,我们提出CIMR:上下文感知的迭代多模态推理框架,包含两个阶段:初始推理与响应生成,随后通过解析的多模态反馈进行迭代优化。动态融合模块在每一步深度融合文本、视觉与上下文特征。我们在Visual Instruction Tuning(VIT)数据集上微调了LLaVA-1.5-7B,并在新提出的多模态动作规划(MAP)数据集上评估CIMR。结果表明,CIMR达到91.5%的准确率,优于GPT-4V(89.2%)、LLaVA-1.5(78.5%)、MiniGPT-4(75.3%)和InstructBLIP(72.8%),验证了其在复杂任务中迭代推理与自我修正的有效性。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) has enhanced our ability to process and generate human language and visual information. However, these models often struggle with complex, multi-step multi-modal instructions that require logical reasoning, dynamic feedback integration, and iterative self-correction. To address this, we propose CIMR: Contextualized Iterative Multimodal Reasoning, a novel framework that introduces a context-aware iterative reasoning and self-correction module. CIMR operates in two stages: initial reasoning and response generation, followed by iterative refinement using parsed multi-modal feedback. A dynamic fusion module deeply integrates textual, visual, and contextual features at each step. We fine-tune LLaVA-1.5-7B on the Visual Instruction Tuning (VIT) dataset and evaluate CIMR on the newly introduced Multi-modal Action Planning (MAP) dataset. CIMR achieves 91.5% accuracy, outperforming state-of-the-art models such as GPT-4V (89.2%), LLaVA-1.5 (78.5%), MiniGPT-4 (75.3%), and InstructBLIP (72.8%), demonstrating the efficacy of its iterative reasoning and self-correction capabilities in complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。