arXiv:2603.07462cs.AI2026-03

用人类感知难度重新定义模型偏差,更准确比较机器与人如何出错。

Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment

  • 以人类识别难易度为基准构建偏差等级谱,替代传统主观设定。
  • 发现不同模型在近距和远距偏差下表现差异:ViT和CNN各有优势。
  • 适合研究模型可信赖性、认知科学或人类对齐的学者参考。

判断人工智能系统是否像人类一样处理信息,是认知科学和可信AI的核心问题。尽管现代AI在标准任务上可达到人类准确率,但这种一致性并不意味着其决策机制与人类相似。通过误差对齐指标比较人类与模型在扭曲或更具挑战性刺激下的失败模式,能更精细地刻画模型与人类的对齐程度。然而,现有分布外(OOD)分析受限于方法选择:要么以模型训练数据为基准定义偏差,要么使用与人类感知无关的任意扭曲参数,难以实现严谨比较。本文提出一种以人为中心的框架,将分布外程度重新定义为人类感知难度的连续谱。通过量化一组刺激相对于无失真参考集的人类准确率下降程度,构建了包含四个明确感知挑战阶段的OOD谱。该方法使模型与人类在可校准难度水平上的比较成为可能。应用于物体识别任务后,揭示了深度学习架构在不同偏差阶段表现出独特的模型-人类对齐排名与特征。视觉语言模型在近距和远距偏差条件下均最接近人类表现;卷积神经网络(CNNs)在近距偏差下比视觉变换器(ViTs)更对齐,而在远距偏差下则相反。本研究强调,在评估模型-人类对齐时,必须考虑跨条件差异(如感知难度)的重要性。

原文摘要 · Abstract (English)

Determining whether AI systems process information similarly to humans is central to cognitive science and trustworthy AI. While modern AI models can match human accuracy on standard tasks, such parity does not guarantee that their underlying decision-making strategies resemble those of humans. Assessing performance using error alignment metrics to compare how humans and models fail, and how this changes for distorted, or otherwise more challenging, stimuli, provides a viable pathway toward a finer characterization of model-human alignment. However, existing out-of-distribution (OOD) analyses for challenging stimuli are limited due to methodological choices: they define OOD shift relative to model training data or use arbitrary distortion-specific parameters with little correspondence to human perception, hindering principled comparisons. We propose a human-centred framework that redefines the degree of OOD as a spectrum of human perceptual difficulty. By quantifying how much a collection of stimuli deviates from an undistorted reference set based on human accuracy, we construct an OOD spectrum and identify four distinct regimes of perceptual challenge. This approach enables principled model-human comparisons at calibrated difficulty levels. We apply this framework to object recognition and reveal unique, regime-dependent model-human alignment rankings and profiles across deep learning architectures. Vision-language models are most consistently human aligned across near- and far-OOD conditions, but convolutional neural networks (CNNs) are more aligned than vision transformers (ViTs) for near-OOD and ViTs are more aligned than CNNs for far-OOD. Our work demonstrates the critical importance of accounting for cross-condition differences, such as perceptual difficulty, for a principled assessment of model-human alignment.

模型对齐认知科学偏差分析视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。