arXiv:2502.16238q-bio.NCcs.AI2025-02被引 22

提出神经AI版图灵测试,要求模型内表征与大脑一致。

Brain-Model Evaluations Need the NeuroAI Turing Test

论文配图:Brain-Model Evaluations Need the NeuroAI Turing Test
图 1 · 摘自论文原文
  • 用行为+内部神经表征双重标准评估模型
  • 模型差异不超过个体大脑间自然差异
  • 适合追求生物真实性的人工智能研究者

什么是好的智能模型?经典图灵测试只关注行为相似性,但不同内部机制可能产生相同输出。在神经人工智能领域,研究常追求模型激活与真实大脑活动的表征一致性。本文指出传统图灵测试对神经人工智能不充分,提出「神经AI图灵测试」新框架:不仅要求行为相似,还要求模型的内部神经表征在可测量个体差异范围内与大脑无法区分——即模型与大脑的差异不大于两个真实大脑间的差异。尽管大脑未必是智能上限,但它是唯一公认的智能实例,适合作为建模基准。该框架推动从模糊的脑启发转向可验证的行为与表征双重标准,为神经科学建模和人工智能发展提供明确评估体系。

原文摘要 · Abstract (English)

What makes an artificial system a good model of intelligence? The classical test proposed by Alan Turing focuses on behavior, requiring that an artificial agent's behavior be indistinguishable from that of a human. While behavioral similarity provides a strong starting point, two systems with very different internal representations can produce the same outputs. Thus, in modeling biological intelligence, the field of NeuroAI often aims to go beyond behavioral similarity and achieve representational convergence between a model's activations and the measured activity of a biological system. This position paper argues that the standard definition of the Turing Test is incomplete for NeuroAI, and proposes a stronger framework called the ``NeuroAI Turing Test'', a benchmark that extends beyond behavior alone and \emph{additionally} requires models to produce internal neural representations that are empirically indistinguishable from those of a brain up to measured individual variability, i.e. the differences between a computational model and the brain is no more than the difference between one brain and another brain. While the brain is not necessarily the ceiling of intelligence, it remains the only universally agreed-upon example, making it a natural reference point for evaluating computational models. By proposing this framework, we aim to shift the discourse from loosely defined notions of brain inspiration to a systematic and testable standard centered on both behavior and internal representations, providing a clear benchmark for neuroscientific modeling and AI development.

神经AI图灵测试表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。