arXiv:2604.00016cs.CLcs.AI2026-04

用人类记忆容量限制检测AI,让大模型暴露真身。

Are they human? Detecting large language models by probing human memory constraints

  • 通过序列回忆任务测试工作记忆,识别机器与人类差异。
  • 即使大模型刻意模仿人类记忆,仍能被认知模型准确区分。
  • 适合在线实验设计者用于防范自动化数据造假。

在线行为研究的有效性依赖于参与者为真实人类而非机器。过去可通过人类易解而机器难解的简单挑战识别机器,但基于大语言模型(LLMs)的通用智能体已能解决多数此类挑战,威胁研究可信度。本文提出新思路:利用机器表现过优反而暴露身份的反向策略。具体地,我们探测人类固有的认知约束——有限的工作记忆容量,在标准序列回忆任务中进行认知建模。结果显示,即便大模型被明确指示模仿人类工作记忆特性,仍可被有效区分。研究表明,利用成熟认知现象可可靠识别大模型与人类。

原文摘要 · Abstract (English)

The validity of online behavioral research relies on study participants being human rather than machine. In the past, it was possible to detect machines by posing simple challenges that were easily solved by humans but not by machines. General-purpose agents based on large language models (LLMs) can now solve many of these challenges, threatening the validity of online behavioral research. Here we explore the idea of detecting humanness by using tasks that machines can solve too well to be human. Specifically, we probe for the existence of an established human cognitive constraint: limited working memory capacity. We show that cognitive modeling on a standard serial recall task can be used to distinguish online participants from LLMs even when the latter are specifically instructed to mimic human working memory constraints. Our results demonstrate that it is viable to use well-established cognitive phenomena to distinguish LLMs from humans.

大模型检测认知建模工作记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。