视觉任务不能靠事后语言推理弥补,必须在视觉空间中实时推理。
Position: Reasoning After Perception Means Reasoning Without Vision
- 将推理从视觉空间移到文本空间,导致关键空间信息丢失
- 实验显示纯文本推理无法解决需在视觉空间完成的任务
- 建议重构架构,让推理直接作用于像素级视觉表示
多模态研究普遍认为,视觉-语言模型的感知缺陷可通过更强的语言推理(如思维链、上下文学习或外部工具)弥补。本文挑战这一假设,指出对于大量难以用语言描述的视觉任务,失败根源在于结构缺陷:何时推理决定了推理发生在何处。当视觉推理被推迟到语言生成阶段,现有架构不仅延迟计算,更将推理从连续视觉表示转移到离散文本空间。因此,传统的‘感知后推理’范式使感知退化为一次性特征编码,功能上等同于‘文本空间中的推理’,导致任务关键的空间信号在推理前已坍缩。我们通过图灵眼测试(TET)验证:必须在视觉空间解决且难以言说的任务,纯文本推理无法修复感知失败。研究建议重新思考架构设计,从‘关于感知的推理’转向‘在感知中推理’,实现直接作用于像素级视觉表示的主动推理驱动感知。
原文摘要 · Abstract (English)
A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-context learning, or external tools). We challenge this assumption. We argue that for a broad class of visual tasks hard to specify in language, failures stem from a structural fatality where the temporal decision of \textit{when} to reason strictly dictates the spatial constraint of \textit{where} reasoning takes place. When visual reasoning is deferred to language generation, current architectures do not merely delay computation; they displace it from the continuous visual representation to a discrete textual space. Consequently, the sequential ``Perception-then-Reasoning'' paradigm degenerates perception into a passive, one-off feature encoding process, rendering it functionally equivalent to ``Reasoning-in-Text-Space'', where task-critical spatial signals are collapsed before reasoning begins. We substantiate this claim with the Turing Eye Test (TET): tasks that must be resolved in \emph{visual space} and are hard to verbalize; results show text-only reasoning cannot remedy these perceptual failures. Our findings suggest rethinking the architectural divide: shifting from reasoning \textit{about} perception to reasoning \textit{within} perception. This facilitates actively reasoning-driven perception that operates directly on pixel-level visual representations, rather than within a collapsed textual space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。