arXiv:2511.19872cs.AI2025-11被引 2

用心理测量法测试大模型自我评估能力,发现其表现稳定但非人类式自信心。

Simulated Self-Assessment in Large Language Models: A Psychometric Approach to AI Self-Efficacy

  • 将人类自效能量表改编为测试大模型自我评估,设计三类任务引导场景。
  • 模型在重复测试中得分高度一致,但整体自评分数低于人类且波动大。
  • 适合关注模型表达能力、局限性与人工系统自我认知的读者。

大型语言模型(LLM)在评估自身能力方面仍不明确。本研究采用10项通用自我效能量表(GSES)对10个主流大模型进行控制性心理测量实验,在无任务控制条件及计算推理、社会推理、摘要三种任务诱导条件下完成测评。分析涵盖模型间差异、任务与无任务对比、模型内稳定性、题序效应、内部一致性及定性推理模式。结果显示,模型生成的GSES回答在多次测试中几乎完全一致,所有模型-任务-题项组合在三次运行中得分不变;内部一致性高(Cronbach's alpha),题序影响极小(组内相关系数)。尽管如此,不同模型在所有条件下自评得分存在显著差异。综合得分低于人类常模(文献对照),潜在响应结构未复现人类自效能模式。定性分析表明,模型对努力、应对、自主性和坚持等概念理解各异:部分模型认为这些属性不适用于人工智能系统,另一些则将其转化为任务能力表述。结果表明,大模型可生成内部一致的心理测量自评,但这类输出反映的是模型特异性和情境敏感的表达行为,而非人类式的自我效能或元认知。心理测量提示法或有助于揭示大模型在不同提示下如何表达能力、局限与自主性。

原文摘要 · Abstract (English)

Large language model (LLM) proficiency in evaluating and quantifying their capacities remains uncertain. We conducted a controlled psychometric measurement study adapting the 10-item General Self-Efficacy Scale (GSES) to evaluate simulated self-assessment across 10 contemporary LLMs. Models completed the GSES under a no-task control condition and after three task-priming conditions: computational reasoning, social reasoning, and summarization. We examined between-model variation, task versus no-task differences, within-model stability, item-order robustness, internal consistency, and qualitative reasoning patterns. LLM-generated GSES responses were highly stable across repeated administrations, with nearly all model-task-item scores remaining identical across three runs. Internal consistency was high across task conditions, with Cronbachs alpha , and item-order effects were minimal, with intraclass correlation coefficients . Despite this stability, simulated self-efficacy differed significantly between models across all task and no-task conditions . Composite LLM GSES scores were lower and more variable than human norms (compared to the literature), and the latent response structure did not replicate human self-efficacy patterns. Qualitative analysis suggested that models differed in how they interpreted GSES items involving effort, coping, agency, and persistence. Some models rejected these constructs as inapplicable to artificial systems, while others translated them into task-oriented capability language. These findings indicate that LLMs can generate internally consistent psychometric self-assessments, but these outputs appear to reflect model-specific and context-sensitive communication behavior rather than human-like self-efficacy or metacognition. Psychometric prompting may help characterize how LLMs express capability, limitation, and agency under different prompting contexts.

大模型自我评估心理测量元认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。