研究大模型间信任的建立与测量,发现显式提问可能误导判断。
Building and Measuring Trust between Large Language Models
- 通过动态建立关系、预设信任脚本、调整系统提示三种方式构建信任
- 显式信任问卷与隐式说服/合作倾向测得结果高度负相关
- 建议用情境化隐式指标评估模型间信任,而非直接询问
随着大语言模型在多智能体系统中日益频繁交互,我们期望它们之间能发展出类似人类同事、朋友或伴侣的信任关系。尽管已有研究表明大模型能识别情感联结并理解信任博弈中的互惠性,但关于(1)不同信任构建策略的比较,(2)如何隐式测量信任,以及(3)隐式与显式信任的关系仍知之甚少。本文通过将隐式信任指标(如易受说服程度、财务合作意愿)与心理学中成熟的人际信任问卷(显式测量)进行对比,探讨上述问题。实验采用三种信任构建方式:动态建立亲和关系、使用预设体现信任的脚本、调整系统提示。结果令人意外:显式信任评分与隐式指标要么弱相关,要么呈显著负相关。这表明,仅通过询问模型对彼此的信任程度可能产生误导。因此,理解模型间信任应更多依赖情境化的隐式测量方法。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly interact with each other, most notably in multi-agent setups, we may expect (and hope) that `trust' relationships develop between them, mirroring trust relationships between human colleagues, friends, or partners. Yet, though prior work has shown LLMs to be capable of identifying emotional connections and recognizing reciprocity in trust games, little remains known about (i) how different strategies to build trust compare, (ii) how such trust can be measured implicitly, and (iii) how this relates to explicit measures of trust. We study these questions by relating implicit measures of trust, i.e. susceptibility to persuasion and propensity to collaborate financially, with explicit measures of trust, i.e. a dyadic trust questionnaire well-established in psychology. We build trust in three ways: by building rapport dynamically, by starting from a prewritten script that evidences trust, and by adapting the LLMs' system prompt. Surprisingly, we find that the measures of explicit trust are either little or highly negatively correlated with implicit trust measures. These findings suggest that measuring trust between LLMs by asking their opinion may be deceiving. Instead, context-specific and implicit measures may be more informative in understanding how LLMs trust each other.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。