arXiv:2411.03630cs.AIq-bio.NC2024-11NeurIPS被引 10

让神经网络决策速度匹配人类反应时间,提升视觉模型真实性。

RTify: Aligning Deep Neural Networks with Human Behavioral Decisions

  • 用RNN模拟人类决策动态,通过反应时约束时间步数。
  • 模型准确拟合人类反应时数据,实现速度与精度平衡。
  • 适合研究人机认知对齐或生物启发的视觉系统设计者。

当前灵长类视觉的神经网络模型多关注整体行为准确率,常忽略感知决策的丰富动态特性。本文提出一种新计算框架,通过学习将循环神经网络(RNN)的时间动态与人类反应时间(RTs)对齐,建模人类行为选择的动态过程。我们设计了一种近似方法,可约束RNN完成任务所需的时间步数,使其与人类反应时间一致。该方法在多个心理物理学实验中得到广泛验证。此外,该近似可用于优化“理想观察者”RNN模型,在无真实人类数据情况下实现速度与准确性的最优权衡。所得模型能良好拟合人类反应时数据。最后,我们利用该近似训练了一个深度学习版的流行Wong-Wang决策模型,并将其与卷积神经网络(CNN)视觉处理模型结合,使用人工及自然图像进行评估。总体而言,本文提出一个新框架,使现有视觉模型更贴近人类行为,推动构建统一的人类视觉模型。

原文摘要 · Abstract (English)

Current neural network models of primate vision focus on replicating overall levels of behavioral accuracy, often neglecting perceptual decisions' rich, dynamic nature. Here, we introduce a novel computational framework to model the dynamics of human behavioral choices by learning to align the temporal dynamics of a recurrent neural network (RNN) to human reaction times (RTs). We describe an approximation that allows us to constrain the number of time steps an RNN takes to solve a task with human RTs. The approach is extensively evaluated against various psychophysics experiments. We also show that the approximation can be used to optimize an "ideal-observer" RNN model to achieve an optimal tradeoff between speed and accuracy without human data. The resulting model is found to account well for human RT data. Finally, we use the approximation to train a deep learning implementation of the popular Wong-Wang decision-making model. The model is integrated with a convolutional neural network (CNN) model of visual processing and evaluated using both artificial and natural image stimuli. Overall, we present a novel framework that helps align current vision models with human behavior, bringing us closer to an integrated model of human vision.

认知建模反应时视觉认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。