让大模型更懂复杂视觉推理,无需额外训练即可提升表现。
Enhancing Advanced Visual Reasoning Ability of Large Language Models
- 用迭代自优化描述图像,结合大模型文本推理能力
- 在多个复杂视觉推理任务上达到当前最优水平
- 适合研究多模态推理与大模型应用的学者
视觉语言(VL)研究的最新进展催生了针对复杂视觉推理的新基准,对模型的高级推理能力提出了挑战。传统视觉语言模型(VLM)在视觉感知任务中表现良好,但在复杂推理场景中表现不佳;而大语言模型(LLM)虽具备强大的文本推理能力,却缺乏视觉理解力。为此,我们提出复杂视觉推理大语言模型(CVR-LLM),充分利用VLM的视觉感知优势与LLM的文本推理能力。不同于需投影层的近期多模态大模型(MLLM),我们的方法通过迭代自优化循环将图像转化为详细、上下文敏感的描述,并直接利用LLM的文本知识进行精准预测,无需额外训练。我们还引入一种新型多模态上下文学习(ICL)方法,以增强LLM的上下文理解与推理能力。此外,提出链式对比(CoC)技术,实现对预测结果各方面的逐步对比分析。CVR-LLM首次系统性地覆盖广泛复杂的视觉推理任务,整体性能达到当前最优(SOTA)。
原文摘要 · Abstract (English)
Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Conversely, Large Language Models (LLMs) demonstrate robust text reasoning capabilities; however, they lack visual acuity. To bridge this gap, we propose Complex Visual Reasoning Large Language Models (CVR-LLM), capitalizing on VLMs' visual perception proficiency and LLMs' extensive reasoning capability. Unlike recent multimodal large language models (MLLMs) that require a projection layer, our approach transforms images into detailed, context-aware descriptions using an iterative self-refinement loop and leverages LLMs' text knowledge for accurate predictions without extra training. We also introduce a novel multi-modal in-context learning (ICL) methodology to enhance LLMs' contextual understanding and reasoning. Additionally, we introduce Chain-of-Comparison (CoC), a step-by-step comparison technique enabling contrasting various aspects of predictions. Our CVR-LLM presents the first comprehensive study across a wide array of complex visual reasoning tasks and achieves SOTA performance among all.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。