arXiv:2508.17298cs.CVcs.AI2025-08综述被引 30

系统梳理视觉推理的组合式发展脉络,助力机器像人一样拆解场景、逻辑推断。

Explain Before You Answer: A Survey on Compositional Visual Reasoning

  • 构建五阶段演化框架,解析从提示到智能体的视觉推理范式变迁。
  • 归纳60+评测基准与指标,覆盖对齐精度、思维链忠实度等核心维度。
  • 揭示幻觉、推理偏见等挑战,指引世界模型融合与人机协同新方向。

组合式视觉推理已成为多模态AI的关键研究前沿,旨在赋予机器人类般的分解视觉场景、定位中间概念及执行多步逻辑推理能力。尽管早期综述聚焦于整体性视觉语言模型或通用多模态推理,但针对快速发展的组合式视觉推理文献尚缺乏专门整合。本文填补这一空白,系统回顾2023至2025年间来自CVPR、ICCV、NeurIPS、ICML、ACL等顶会的260余篇论文。首先形式化核心定义,阐明组合方法在认知契合度、语义保真度、鲁棒性、可解释性与数据效率方面的优势。接着追溯五阶段范式演进:从提示增强的语言中心流程,经工具增强的大语言模型与视觉语言模型,到近期兴起的思维链推理与统一智能体式视觉语言模型,剖析其架构设计、优势与局限。随后整理60多个基准与对应指标,用于评估组合视觉推理在定位精度、思维链忠实度及高分辨率感知等方面的表现。基于上述分析,提炼关键洞见,识别开放挑战(如基于大语言模型推理的局限、幻觉、演绎推理偏好、可扩展监督、工具集成与基准缺陷),并提出未来方向,包括世界模型融合、人机协同推理与更丰富的评估协议。本综述提供统一分类体系、历史路线图与批判性展望,旨在成为该领域基础参考,并激发下一代组合式视觉推理研究。

原文摘要 · Abstract (English)

Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground intermediate concepts, and perform multi-step logical inference. While early surveys focus on monolithic vision-language models or general multimodal reasoning, a dedicated synthesis of the rapidly expanding compositional visual reasoning literature is still missing. We fill this gap with a comprehensive survey spanning 2023 to 2025 that systematically reviews 260+ papers from top venues (CVPR, ICCV, NeurIPS, ICML, ACL, etc.). We first formalize core definitions and describe why compositional approaches offer advantages in cognitive alignment, semantic fidelity, robustness, interpretability, and data efficiency. Next, we trace a five-stage paradigm shift: from prompt-enhanced language-centric pipelines, through tool-enhanced LLMs and tool-enhanced VLMs, to recently minted chain-of-thought reasoning and unified agentic VLMs, highlighting their architectural designs, strengths, and limitations. We then catalog 60+ benchmarks and corresponding metrics that probe compositional visual reasoning along dimensions such as grounding accuracy, chain-of-thought faithfulness, and high-resolution perception. Drawing on these analyses, we distill key insights, identify open challenges (e.g., limitations of LLM-based reasoning, hallucination, a bias toward deductive reasoning, scalable supervision, tool integration, and benchmark limitations), and outline future directions, including world-model integration, human-AI collaborative reasoning, and richer evaluation protocols. By offering a unified taxonomy, historical roadmap, and critical outlook, this survey aims to serve as a foundational reference and inspire the next generation of compositional visual reasoning research.

视觉推理多模态大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。