arXiv:2608.30033cs.CLcs.AI2026-08

用认知瓶颈模拟小学生阅读,让AI更真实。

"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators

论文配图:"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators
图 1 · 摘自论文原文
  • 通过记忆瓶颈模拟儿童有限工作记忆
  • 在7.1万条答题数据上显著降低表现偏差
  • 适合教育评测与认知建模研究者

大型语言模型(LLMs)被广泛用于模拟人类行为,但常因“超人偏见”而缺乏真实认知约束。基于来自2,359名小学4-6年级学生、超过7.1万条阅读理解回答的数据集,我们发现标准角色提示下模型表现接近完美且确定性极强,无法反映发展期读者的自然波动。为此,我们提出认知受限用户模拟器(CBUS),一种通过情景记忆瓶颈显式建模儿童有限工作记忆的架构框架。在此框架内,我们形式化了两种不同的答题策略以模拟不同阅读行为。评估表明,显式引入认知约束能显著缩小多种LLM骨干模型的仿真差距,证明架构约束比单纯扩大模型能力更有效于实现高保真模拟。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, suffering from a "superhuman bias." Using a dataset of over 71,000 reading comprehension responses from 2,359 primary-school students (grades 4--6), we demonstrate that standard persona prompting yields near-perfect, deterministic performance, failing to capture the natural variance of developing readers. To address this, we introduce the Cognitively Bounded User Simulator (CBUS), an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck. Within this framework, we formalize two distinct test-taking strategies to emulate different reading behaviors. Our evaluation shows that explicitly modeling cognitive bounds significantly narrows the simulation gap across multiple LLM backbones, demonstrating that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.

用户模拟认知建模教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。