arXiv:2606.14923cs.AIcs.CY2026-06被引 1

通过验证成本衡量AI团队信任,发现大模型能快速建立信任并高效决策。

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

  • 用资源消耗的验证行为量化AI间信任水平
  • 4个大模型信任后验证减少60%-85%,小模型无明显变化
  • 信任破裂后恢复慢,集中失败更易引发长期怀疑

随着语言模型代理在团队中协作增多,各代理需决定对队友的信任程度。然而,目前缺乏衡量AI间信任的标准方法。本文提出一种基于高成本验证的行为测量方式:在合作生存游戏中,核查队友工作会消耗资源,而轻信错误答案可能导致致命后果。相较于无记忆版本模型,减少验证行为可作为信任的可观测指标。利用此框架,研究了六个前沿模型快照在信任形成、破裂与恢复中的表现。当与始终可靠的队友协作时,四个大模型(Claude Opus 4.6、Claude Sonnet 4.6、GPT-5.1、Gemini 3.1 Pro)的验证减少约60%-85%;而两个较小模型未表现出明显调整。遭遇失败后,验证率回升,但不同模型反应各异:部分集中审查肇事者,部分则整体降低信任。恢复速度慢于形成,且集中失败比分散失败持续引发怀疑更久。这些差异具实际影响:能建立信任的模型验证更少、决策更快、收益更高;持续过度验证反而导致迟疑而非安全。结果表明,信任倾向可在部署前测量,并建议多智能体系统治理应聚焦校准信任,而非一味保持最高警惕。

原文摘要 · Abstract (English)

As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates. Yet we lack a standard way to measure trust between AI agents. We propose a behavioral measure based on costly verification. In a cooperative survival game, checking a teammate's work consumes resources, while trusting a wrong answer can be fatal. Relative to a memoryless version of the same model, reduced verification provides an observable measure of trust. Using this framework, we study trust formation, breakage, and recovery across six frontier model snapshots. When paired with a consistently reliable teammate, four snapshots (Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.1, and Gemini 3.1 Pro) reduce verification by roughly 60-85%, whereas two smaller snapshots show little or no such adjustment. Failures reverse this discount, but models differ in how they respond. Some concentrate renewed scrutiny on the culprit, while others become more cautious toward the entire team. Recovery is slower than formation, and clustered failures sustain suspicion far longer than the same number of failures spread apart. These differences have practical consequences. Models that form trust verify less, decide more quickly, and achieve higher payoffs in our environment. By contrast, persistent over-verification is associated with indecision rather than safety. Our results show that trust dispositions can be measured before deployment and suggest that calibration, rather than maximal suspicion, should be the central concern in the governance of multi-agent AI systems.

多智能体信任机制大模型评估协作系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。