arXiv:2505.20164cs.CL2025-05被引 7

用视觉抽象提升多模态模型推理能力,更高效准确。

Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction

  • 用视觉摘要替代复杂文字推理,聚焦关键视觉特征。
  • 相比传统方法平均提升2.21%性能,且节省计算资源。
  • 适合追求高效多模态推理的开发者与研究者。

图像通常比文本包含更丰富的细节,但常带有冗余信息,可能降低多模态推理性能。人类在面对复杂信息时,倾向于通过抽象思维将其简化为简洁的抽象表达。受此认知策略启发,我们提出一种新范式——视觉抽象思考(VAT),通过向多模态大语言模型(MLLMs)输入视觉摘要,而非显式文字思考或详细引导,实现更高效的视觉推理机制。VAT通过削弱冗余信息,使模型更关注本质视觉元素、概念和结构特征,相比链式思维(CoT)和工具使用等方法,避免了冗长中间步骤和外部知识引入带来的复杂性。实验表明,VAT在多种视觉感知与推理任务中持续提升不同MLLM的表现,平均超越GPT-5基线2.21%,优于CoT效果,并在更低的令牌消耗下实现更高性能。结果验证了视觉抽象思维的有效性,也推动从人类认知角度探索更多元的推理范式。

原文摘要 · Abstract (English)

Images usually convey richer detail than text, but often include redundant information, which potentially downgrades multimodal reasoning performance. When faced with lengthy or complex messages, humans tend to employ abstract thinking to convert them into simple and concise abstracts. Inspired by this cognitive strategy, we introduce a novel paradigm to elicit the ability to Think with Visual Abstract (VAT), by prompting Multimodal Large Language Models (MLLMs) with visual abstract instead of explicit verbal thoughts or elaborate guidance, permitting a more efficient visual reasoning mechanism via concentrated perception. VAT encourages models to focus on more essential visual elements, concepts and structural features by undermining redundant information compared with explicit thinking methods, such as Chain-of-thought (CoT) and tool-using approaches, that increase the complexity of reasoning process via inserting verbose intermediate steps and external knowledge. Experimental results show that VAT consistently empowers different MLLMs in visual perception and reasoning tasks. VAT achieves an average gain of $2.21\%$ over GPT-5 baseline, surpassing the gain of CoT, demonstrating that VAT better enhances multimodal task performance of MLLMs. Additionally, VAT spends fewer tokens while achieving higher performance. These findings highlight the effectiveness of visual abstract thinking and encourage further exploration of more diverse reasoning paradigms from the perspective of human cognition.

多模态推理视觉抽象MLLM高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。