arXiv:2603.24742cs.AIcs.LG2026-03被引 1

研究用户如何在反复交互中动态建立对AI的信任,揭示安全与监管的关键条件。

Trust or Check? Understanding the (Evolutionary) Dynamics of User Trust in AI Systems

  • 用演化博弈论建模用户与AI开发者间的持续互动信任机制。
  • 发现三种长期状态:无人使用、不安全广泛使用、安全广泛使用。
  • 只有当惩罚成本高于安全成本且用户能偶尔监督时,安全系统才能普及。

随着人工智能系统能力与应用的提升,用户对其信任问题日益突出。现有研究多聚焦于治理模型,将用户信任视为一次性采纳决策,而非由重复互动塑造的动态过程。本文将信任建模为用户在用户与开发者不对称互动中减少监控的动态选择,其中检查开发者行为具有成本。基于演化博弈论,分析在不同监控成本和制度环境下,用户信任策略与开发者提供安全或不安全AI的策略如何共演化。采用基于模仿与学习的视角,结合随机有限种群动态、无限种群复制者动力学及强化学习分析。结果表明存在三种稳健的长期状态:用户不采纳但开发者提供不安全系统;不安全但广泛采用;安全且广泛采用。仅最后一种理想,其出现需满足:对不安全行为的惩罚超过安全成本,且用户仍能偶尔进行监控。研究支持强调透明度、低成本监控和有效惩戒的治理方案,并表明仅靠监管或盲目信任均不足以避免不安全或低采纳率的结果。

原文摘要 · Abstract (English)

As the capabilities and adoption of Artificial Intelligence (AI) systems grow, trust in these AI systems is an increasingly urgent concern. Much research has focused on models of AI governance and has primarily examined incentives for safe development and effective regulation. Hence they typically represented users trust as a one-shot adoption choice rather than as a dynamic, evolving process shaped by repeated interactions. We instead model trust as the dynamic choice of reduced monitoring in a repeated, asymmetric interaction between users and AI developers, where checking developers' behaviour is costly. Using evolutionary game theory, we study how users' strategies of trust and developers' strategies of providing safe (compliant) or unsafe (non-compliant) AI co-evolve under different levels of monitoring cost and institutional regimes. We conduct the analysis on both imitation-based and learning-based perspectives, with the stochastic finite-population dynamics, the infinite-population replicator analysis and the reinforcement learning analysis. We find three robust long-run regimes: no adoption by users while developers provide unsafe AI, unsafe but widely adopted systems, and safe systems that are widely adopted. Only the last is desirable, and it arises when penalties for unsafe behaviour exceed the extra cost of safety and users can still afford to monitor at least occasionally. Our results formally support governance proposals that emphasise transparency, low-cost monitoring, and meaningful sanctions, and they show that neither regulation alone nor blind user trust is sufficient to prevent the drift towards unsafe or low-adoption outcomes.

AI信任演化博弈治理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。