arXiv:2602.00769cs.CLcs.AI2026-02

用博弈实验测大模型信任度,发现GPT-4.1最像人。

Eliciting Trustworthiness Priors of Large Language Models via Economic Games

  • 用信任博弈游戏诱导模型暴露信任倾向
  • GPT-4.1的信任行为与人类高度一致
  • 模型能区分不同角色的可信度,受暖感和能力感知影响

构建以人为本、可信赖的人工智能系统,关键在于保持校准的信任——既不过度依赖(如自动化偏见),也不过度怀疑(如弃用)。但如何刻画AI自身表现出的信任水平仍是一大挑战。本文提出一种基于迭代上下文学习的新方法,应用于行为博弈论中的信任博弈(Trust Game),以量化大语言模型的可信度先验。该博弈将信任定义为基于对他人信念的自愿风险承担,而非主观报告。我们对多个主流大模型进行测试,发现GPT-4.1的可信度先验与人类表现高度匹配。进一步分析显示,该模型能根据玩家人格特征差异化回应信任行为。最后,我们验证了所提取的信任差异可被基于刻板印象的模型有效预测,其核心依据是感知到的温暖度与能力感。

原文摘要 · Abstract (English)

One critical aspect of building human-centered, trustworthy artificial intelligence (AI) systems is maintaining calibrated trust: appropriate reliance on AI systems outperforms both overtrust (e.g., automation bias) and undertrust (e.g., disuse). A fundamental challenge, however, is how to characterize the level of trust exhibited by an AI system itself. Here, we propose a novel elicitation method based on iterated in-context learning (Zhu and Griffiths, 2024a) and apply it to elicit trustworthiness priors using the Trust Game from behavioral game theory. The Trust Game is particularly well suited for this purpose because it operationalizes trust as voluntary exposure to risk based on beliefs about another agent, rather than self-reported attitudes. Using our method, we elicit trustworthiness priors from several leading large language models (LLMs) and find that GPT-4.1's trustworthiness priors closely track those observed in humans. Building on this result, we further examine how GPT-4.1 responds to different player personas in the Trust Game, providing an initial characterization of how such models differentiate trust across agent characteristics. Finally, we show that variation in elicited trustworthiness can be well predicted by a stereotype-based model grounded in perceived warmth and competence.

大模型信任博弈实验可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。