用心理学框架检验大模型认知行为,发现其表现类人但受训练数据影响。
AI Through the Human Lens: Investigating Cognitive Theories in Machine Psychology
- 通过心理测试框架评估大模型的叙事与决策模式。
- 模型对积极表述敏感,道德判断偏向自由/压迫议题,常自我矛盾。
- 揭示大模型认知类人特性,适合关注AI伦理与安全的研究者阅读。
我们基于心理学四大理论框架——主题统觉测验(TAT)、框架效应、道德基础理论(MFT)和认知失调——探究大型语言模型(LLMs)是否表现出类人认知特征。采用结构化提示与自动评分方法,评估了多个专有及开源模型。结果表明,这些模型常生成连贯叙事,易受正面表述影响,道德判断倾向于自由/压迫议题,并表现出由大量理性化缓解的自我矛盾。此类行为虽与人类认知倾向相似,却由训练数据和对齐机制塑造。研究讨论了对AI透明度、伦理部署及认知心理学与AI安全融合的启示。
原文摘要 · Abstract (English)
We investigate whether Large Language Models (LLMs) exhibit human-like cognitive patterns under four established frameworks from psychology: Thematic Apperception Test (TAT), Framing Bias, Moral Foundations Theory (MFT), and Cognitive Dissonance. We evaluated several proprietary and open-source models using structured prompts and automated scoring. Our findings reveal that these models often produce coherent narratives, show susceptibility to positive framing, exhibit moral judgments aligned with Liberty/Oppression concerns, and demonstrate self-contradictions tempered by extensive rationalization. Such behaviors mirror human cognitive tendencies yet are shaped by their training data and alignment methods. We discuss the implications for AI transparency, ethical deployment, and future work that bridges cognitive psychology and AI safety
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。