通过社会世界观框架,揭示大模型在权威、平等等方面的隐性认知差异。
Analyzing Cognitive Differences Among Large Language Models through the Lens of Social Worldview
- 构建社会世界观分类体系,量化模型对权威、平等、自主和命运的认知倾向。
- 分析28个大模型发现其具有独特社会认知特征,且可被社会提示动态调节。
- 为开发更透明、负责任的AI提供可解释的认知评估工具,适合伦理与AI安全研究者。
大语言模型深刻影响社会互动、决策与信息传播,亟需理解其隐含的社会认知态度,即“世界观”。不同于以往将人口与伦理偏见视为固定属性的研究,本文聚焦权威、平等、自主与命运等深层认知取向,强调其在动态社会情境中的可塑性。我们提出社会世界观分类框架(SWT),基于文化理论将四种经典世界观——等级制、平等主义、个人主义与宿命论——转化为可量化的子维度。通过对28个不同大模型的广泛分析,识别出反映内在社会认知结构的独特认知画像。借助社会参照理论,实验表明显式社会线索能系统性调节这些画像,揭示出模型强大的认知适应能力。研究为理解大模型的潜在认知灵活性提供了洞见,并为计算科学家开发更透明、可解释、负责任的AI系统提供了实用路径。
原文摘要 · Abstract (English)
Large Language Models significantly influence social interactions, decision-making, and information dissemination, underscoring the need to understand the implicit socio-cognitive attitudes, referred to as "worldviews", encoded within these systems. Unlike previous studies predominantly addressing demographic and ethical biases as fixed attributes, our study explores deeper cognitive orientations toward authority, equality, autonomy, and fate, emphasizing their adaptability in dynamic social contexts. We introduce the Social Worldview Taxonomy (SWT), an evaluation framework grounded in Cultural Theory, operationalizing four canonical worldviews, namely Hierarchy, Egalitarianism, Individualism, and Fatalism, into quantifiable sub-dimensions. Through extensive analysis of 28 diverse LLMs, we identify distinct cognitive profiles reflecting intrinsic model-specific socio-cognitive structures. Leveraging principles from Social Referencing Theory, our experiments demonstrate that explicit social cues systematically modulate these profiles, revealing robust patterns of cognitive adaptability. Our findings provide insights into the latent cognitive flexibility of LLMs and offer computational scientists practical pathways toward developing more transparent, interpretable, and socially responsible AI systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。