arXiv:2605.08549cs.AI2026-05被引 1

用20题测试评估大模型在认知发展层面的回应模式。

Evaluating Developmental Cognition Capabilities of LLMs

论文配图:Evaluating Developmental Cognition Capabilities of LLMs
图 1 · 摘自论文原文
  • 设计20项自填式句子补全测试,捕捉认知发展阶段信号。
  • 真实人类回答中人与模型一致性为中等,同类回答趋同性强。
  • 大模型生成文本呈现稳定发展阶段差异,越新越大模型表现更高阶。

对话式AI日益根据用户偏好、经历、目标和知识进行个性化,但较少关注用户如何解读和运用模型输出来构建与理解自身现实。本文借鉴罗伯特·基根的建构-发展理论,提出一种名为发展性句子补全测试(DSCT)的20项自填工具,用于在无需专家访谈的前提下,从文本中提取认知发展信号。我们不将标签视为个体发展阶段的验证结果,而是将其作为响应中阶段结构的表征。研究考察了大模型在三种响应情境下的表现:模拟人格、真实人类回答及默认模型生成。在模拟人格下,前沿模型对意图标签的识别准确率高;在真实人类数据上,人机一致性为中等,同类反应间一致性显著高于完全匹配;当模型无角色条件生成时,不同模型家族展现出稳定的阶段特征,更大更先进的模型生成的文本评分更高。结果表明,合成数据中的发展信号更清晰,而实现阶段感知型对话的核心挑战并非分类器精度,而是获取可解析的发展信号。

原文摘要 · Abstract (English)

Conversational AI is increasingly personalized around users' preferences, histories, goals, and knowledge, but much less around how users interpret and take up model outputs to construct and understand their reality. We draw on Robert Kegan's constructive-developmental theory as a complementary lens on this dimension. Existing methods for assessing developmental stage in the Keganian tradition rely either on expert interviews that do not scale or on sentence-completion instruments that are proprietary, lengthy, or invasive. To make this perspective tractable for LLM evaluation, we introduce the Developmental Sentence Completion Test (DSCT), a 20-item instrument designed to elicit developmental signal in self-administered text. Throughout, we treat the resulting labels as characterizations of stage-like structure in elicited responses, not as validated person-level developmental stage. We then ask how much of that signal can be recovered by LLMs across three elicited response regimes: simulated personas, real human respondents, and default model-generated answers. On simulated personas, top frontier models recover simulator-intended labels with high accuracy. On real human DSCT responses, human-LLM agreement is fair, with much stronger within-neighborhood than exact agreement. Finally, when LLMs answer DSCT prompts without persona-conditioning, their responses exhibit stable stage-like differences across model families, with larger and newer models tending to generate higher-rated text. These results suggest that stage-conditioned signal is cleaner in synthetic responses than in human-written DSCT text, and that the core constraint for stage-aware conversational AI is not classifier accuracy alone, but the availability of developmental signal from elicited text.

认知模型大模型评估发展心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。