arXiv:2410.14138cs.CVcs.AI2024-10EMNLP被引 11

提出分阶段视觉推理框架,让模型先主动看图再思考,性能提升13.2%。

ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom

  • 将视觉推理拆为主动感知与文本推理两阶段,解耦能力
  • 在多个基准上平均提升13.2%,闭源与开源模型均有效
  • 可无缝接入大语言模型,生成高质量推理数据

大型视觉语言模型(LVLMs)在视觉理解任务中取得显著进展,但在视觉推理任务中常过度依赖语言知识而忽略图像信息,导致性能下降。针对此问题,我们首先识别现有方法的不足:多模态推理能力有限,且视觉描述不充分或无关。为此,我们将视觉推理过程分解为两个阶段:主动视觉感知(即‘眼睛’)和文本推理(即‘智慧’),提出新型视觉推理框架ProReason。该框架具备解耦的视觉-推理能力与多轮主动感知机制。给定多模态问题时,ProReason 迭代进行主动信息收集与推理,直至得出结论并生成必要且充分的视觉描述。其能力解耦特性支持无缝集成现有大语言模型(LLMs),以弥补LVLMs的推理短板。大量实验表明,ProReason 在多种开源自研与闭源模型的基准测试中均优于现有多步推理框架,平均性能提升达13.2%。此外,引入LLMs使ProReason能生成高质量视觉推理数据,由此训练的蒸馏模型(ProReason-VL 和 ProReason-Q3)在下游任务中表现更优。本研究对现有方案的洞察及解耦视角,为未来基于大语言模型的视觉推理技术提供了新方向。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation. To tackle this issue, we first identify the drawbacks of existing solutions (i.e., limited multi-modal reasoning capacities, and insufficient and irrelevant visual descriptions). We then decompose visual reasoning process into two stages: proactive visual perception (i.e., eyesight) and textual reasoning (i.e., wisdom), and introduce a novel visual reasoning framework named ProReason. This framework features decoupled vision-reasoning capabilities and multi-run proactive perception. Briefly, given a multi-modal question, ProReason iterates proactive information collection and reasoning until the answer can be concluded with necessary and sufficient visual descriptions. Notably, the disassociation of capabilities allows seamless integration of existing large language models (LLMs) to compensate for the reasoning deficits of LVLMs. Our extensive experiments demonstrate that ProReason outperforms existing multi-step reasoning frameworks on various benchmarks for both open-source and closed-source models, with the average performance gain reaching 13.2%. Besides, the integration of LLMs allows ProReason to produce high-quality visual reasoning data, which empowers ProReason-distilled models (i.e., ProReason-VL and ProReason-Q3) to achieve superior performance in downstream tasks. Our insights into existing solutions and the decoupled perspective for feasible integration of LLMs illuminate future research on visual reasoning techniques, especially LLM-assisted ones.

视觉推理多模态大模型解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。