arXiv:2606.29416cs.CVcs.AI2026-06被引 1

测试视觉模型能否在无局部线索时识别物体,发现其能力存在根本性极限。

Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances

论文配图:Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances
图 1 · 摘自论文原文
  • 用句法距离衡量类别可分性,构建全局语义差异但局部无差别的图像任务。
  • 当图像尺度超过临界点后,模型准确率骤降至随机水平,且无法恢复。
  • 揭示当前模型在全局概念理解上的结构性局限,适合关注AI认知边界的研究者。

视觉模型是否真能‘看见’物体,还是仅依赖表面视觉线索?受维特根斯坦观点启发,我们认为模型的识别能力受限于其习得的描述系统。当前视觉模型多通过学习特征表示利用局部统计线索。因此,我们探究当局部线索无法提供稳定区分时,模型是否仍能正确分类。为此,提出句法距离概念,通过映射一个类别到另一个类别的操作对称性来衡量类别可分性:正值暴露可利用的局部特征,零值则要求依赖全局语义而非局部规则。构建了最大方差二值噪声下的视觉自指任务:正样本含闭合正方形,负样本为其余相同但边界像素翻转的正方形。两类在全局语义上不同,但句法距离为零,使局部统计捷径失效。对ResNets和视觉变换器的实验显示一致的相变现象:一旦图像尺度超过临界点,准确率即塌缩至随机猜测水平且未恢复。更大的训练集与模型仅延迟该崩溃,而全局注意力的ViTs反而更早达到极限。结果揭示当前架构在全局概念任务上的结构性能力边界,暗示通用智能可能需要创造新语言,而非复用现有语言。

原文摘要 · Abstract (English)

Can a vision model truly see an object, or does it only fit surface-level visual cues? Following Wittgenstein's view that the limits of language are the limits of the world, we view a model's recognition ability as bounded by the descriptive system it has learned. In current vision models, this system is often realized through learned feature representations that exploit local statistical cues. We therefore ask whether a model can still classify correctly when such local cues provide no stable basis for distinction. We formalize this question with syntactic distance, which measures class separability through the symmetry of the operations mapping one class to the other: positive distance exposes exploitable local features, whereas zero distance requires global semantics rather than local rules. We construct a visual self-referential task in maximum-variance binary noise: positive samples contain a closed square, while negative samples contain an otherwise identical square with one flipped boundary pixel. The two classes differ in global semantics but have zero syntactic distance, making local statistical shortcuts unreliable. Experiments on ResNets and Vision Transformers reveal a consistent phase-transition phenomenon, with accuracy collapsing to random guessing once the image scale crosses a critical point and does not recover within the tested range. Larger training sets and models only delay this collapse, while globally attentive ViTs reach it earlier. These results reveal a structural capability boundary of current architectures on global-concept tasks, suggesting that general intelligence may require creating new language, not reusing an existing one.

视觉理解认知边界句法距离模型局限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。