arXiv:2508.10523cs.CV2025-08被引 25

系统梳理视觉推理的五大类型与方法,推动计算机视觉向深层理解演进。

Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies

  • 按关系、符号、时间、因果、常识五类划分视觉推理,构建统一框架。
  • 涵盖图模型、神经符号系统到多模态大模型的多种实现路径。
  • 揭示评估局限并提出可泛化、可解释的可信视觉系统研究方向。

视觉推理对超越表层物体检测与分类的计算机视觉任务至关重要。尽管在关系、符号、时间、因果和常识推理方面已有进展,现有综述通常仅覆盖单一方向,如视觉问答、场景图生成、神经符号人工智能或多模态思维链,且很少同时分析推理类型、方法与评估协议。本文填补这一空白:通过结构化文献综述,将视觉推理分为五类(关系、符号、时间、因果、常识),并考察从图模型、记忆网络、注意力机制到神经符号系统、视觉语言模型(VLMs)和多模态大语言模型(MLLMs)的实现方式,包括视觉思维链、视觉编程、工具增强及测试时推理。进一步回顾功能正确性、结构一致性和因果有效性等评估协议,并分析其在泛化性、可复现性、忠实度和解释力方面的局限。识别出关键挑战:复杂场景扩展、符号与神经范式深度融合、缺乏全面基准、基础模型的语言先验偏倚与幻觉,以及弱监督下的推理。最后提出研究议程,主张感知与推理结合是实现透明、可信、跨领域模型的关键,尤其在自动驾驶与医学诊断等高风险场景中。

原文摘要 · Abstract (English)

Visual reasoning matters for many computer vision tasks that go beyond surface-level object detection and classification. Despite progress in relational, symbolic, temporal, causal, and commonsense reasoning, existing surveys typically cover only one part of the problem, such as visual question answering, scene-graph generation, neuro-symbolic AI, or multimodal chain-of-thought, and rarely analyze reasoning types, methodologies, and evaluation protocols together. This survey addresses that gap. Following a structured literature review, we group visual reasoning into five major types (relational, symbolic, temporal, causal, and commonsense) and examine how each is implemented across methods that range from graph-based models, memory networks, attention mechanisms, and neuro-symbolic systems to reasoning with vision-language models (VLMs) and multimodal large language models (MLLMs), including visual chain-of-thought, visual programming, and tool-augmented and test-time reasoning. We then review evaluation protocols for functional correctness, structural consistency, and causal validity, and we analyze their limits in generalizability, reproducibility, faithfulness, and explanatory power. We also identify open challenges: scaling to complex scenes, integrating symbolic and neural paradigms more deeply, the shortage of comprehensive benchmarks, language-prior shortcuts and hallucination in foundation models, and reasoning under weak supervision. Finally, we set out a research agenda for vision systems and argue that connecting perception and reasoning is necessary for transparent, trustworthy, and cross-domain models, especially in high-stakes settings such as autonomous driving and medical diagnostics.

视觉推理多模态认知建模评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。