arXiv:2603.05839cs.MAcs.AI2026-03

用对比提示法分析大模型如何理解信任,发现其内部表征最贴近人类社会认知模型。

Evaluating LLM Alignment With Human Trust Models

  • 通过对比提示生成嵌入向量,白盒分析大模型对信任的内部表征。
  • 模型内部信任表征与Castelfranchi模型相似度最高,其次为Marsh模型。
  • 结果可支持人机协作系统设计,适合关注可信AI的研究者。

信任在人类互动与多智能体系统中对有效合作、降低不确定性及引导决策至关重要。尽管如此,对大语言模型(LLMs)如何内部构建和推理信任的理解仍有限。本文对EleutherAI/gpt-j-6B进行白盒分析,采用对比提示法在模型激活空间中生成二元信任及相关人际属性的嵌入向量。首先从五个成熟的人类信任模型中提取信任相关概念;然后通过计算60个通用情感概念间的余弦相似度确定显著对齐阈值;最后测量模型内部信任表征与提取概念间的余弦相似度。结果显示,gpt-j-6B的内部信任表征与Castelfranchi社会认知模型最为接近,其次为Marsh模型。这表明大模型在激活空间中编码了社会认知结构,支持有意义的比较分析,有助于社会认知理论发展,并为人类-AI协作系统设计提供依据。

原文摘要 · Abstract (English)

Trust plays a pivotal role in enabling effective cooperation, reducing uncertainty, and guiding decision-making in both human interactions and multi-agent systems. Although it is significant, there is limited understanding of how large language models (LLMs) internally conceptualize and reason about trust. This work presents a white-box analysis of trust representation in EleutherAI/gpt-j-6B, using contrastive prompting to generate embedding vectors within the activation space of the LLM for diadic trust and related interpersonal relationship attributes. We first identified trust-related concepts from five established human trust models. We then determined a threshold for significant conceptual alignment by computing pairwise cosine similarities across 60 general emotional concepts. Then we measured the cosine similarities between the LLM's internal representation of trust and the derived trust-related concepts. Our results show that the internal trust representation of EleutherAI/gpt-j-6B aligns most closely with the Castelfranchi socio-cognitive model, followed by the Marsh Model. These findings indicate that LLMs encode socio-cognitive constructs in their activation space in ways that support meaningful comparative analyses, inform theories of social cognition, and support the design of human-AI collaborative systems.

信任建模大模型分析社会认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。