arXiv:2512.09340cs.AIcs.CV2025-12

对比人与模型对模糊图像的分类策略,揭示认知差异。

Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration

  • 结合认知科学模型分析人类如何用类比、形状判断等策略决策
  • 发现人类在低分辨率图像上依赖信心调节,而模型仅靠特征提取
  • 适合关注可解释AI与认知对齐的研究者

理解人类与人工智能系统如何解读模糊视觉刺激,有助于揭示感知、推理与决策的本质。本文研究了人在低分辨率、感知退化图像上的标注表现,与深度神经网络进行对比。基于计算认知科学、认知架构与联结主义-符号混合模型,我们分析人类采用类比推理、基于形状的识别及信心调节等策略,与模型的特征驱动处理形成对照。依据Marr三层次假说、西蒙有限理性及塔加德的表征与情感框架,将参与者反应与模型注意力(Grad-CAM)关联分析。人类行为通过ACT-R与Soar认知模型建模,揭示在不确定性下分层且启发式的决策机制。研究发现生物与人工系统在表征、推断与置信校准方面存在关键共性与差异。结果推动未来融合结构化符号推理与联结主义表示的神经符号架构发展,这类架构需遵循具身性、可解释性与认知对齐原则,使AI不仅高效,更可解释且具认知基础。

原文摘要 · Abstract (English)

Understanding how humans and AI systems interpret ambiguous visual stimuli offers critical insight into the nature of perception, reasoning, and decision-making. This paper examines image labeling performance across human participants and deep neural networks, focusing on low-resolution, perceptually degraded stimuli. Drawing from computational cognitive science, cognitive architectures, and connectionist-symbolic hybrid models, we contrast human strategies such as analogical reasoning, shape-based recognition, and confidence modulation with AI's feature-based processing. Grounded in Marr's tri-level hypothesis, Simon's bounded rationality, and Thagard's frameworks of representation and emotion, we analyze participant responses in relation to Grad-CAM visualizations of model attention. Human behavior is further interpreted through cognitive principles modeled in ACT-R and Soar, revealing layered and heuristic decision strategies under uncertainty. Our findings highlight key parallels and divergences between biological and artificial systems in representation, inference, and confidence calibration. The analysis motivates future neuro-symbolic architectures that unify structured symbolic reasoning with connectionist representations. Such architectures, informed by principles of embodiment, explainability, and cognitive alignment, offer a path toward AI systems that are not only performant but also interpretable and cognitively grounded.

认知科学可解释AI神经符号图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。