arXiv:2504.08640cs.AIcs.CY2025-04被引 12

LLM代理在博弈中更不信任,需建立用户与监管者间的良性互动。

Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents

  • 用演化博弈论建模开发者、监管者与用户三方策略选择。
  • 用户完全信任时监管激励有效,但有条件信任会破坏社会契约。
  • 不同LLM测试结果差异大,提示监管系统需考虑模型特性。

目前普遍认为,在人工智能开发生态中建立信任与合作至关重要,以促进可信AI系统的采用。本文通过将大型语言模型(LLM)代理嵌入演化博弈论框架,研究了开发者、监管者与用户之间的复杂互动,模拟在不同监管情境下的战略决策。演化博弈论用于量化各主体面临的困境,而LLM则引入更多复杂性与细微差别,支持重复博弈及人格特质建模。研究发现,战略型AI代理表现出比纯博弈论代理更‘悲观’(不信任且背叛)的倾向。当用户完全信任时,监管激励能有效推动合规;但若信任有条件,则可能破坏‘社会契约’。因此,建立用户信任与监管声誉之间的正向反馈机制,似乎是促使开发者构建安全AI的关键。然而,这种信任的形成程度可能取决于所使用的具体LLM。研究结果为人工智能监管体系提供指导,并有助于预测战略型LLM代理在监管辅助中的表现。

原文摘要 · Abstract (English)

There is general agreement that fostering trust and cooperation within the AI development ecosystem is essential to promote the adoption of trustworthy AI systems. By embedding Large Language Model (LLM) agents within an evolutionary game-theoretic framework, this paper investigates the complex interplay between AI developers, regulators and users, modelling their strategic choices under different regulatory scenarios. Evolutionary game theory (EGT) is used to quantitatively model the dilemmas faced by each actor, and LLMs provide additional degrees of complexity and nuances and enable repeated games and incorporation of personality traits. Our research identifies emerging behaviours of strategic AI agents, which tend to adopt more "pessimistic" (not trusting and defective) stances than pure game-theoretic agents. We observe that, in case of full trust by users, incentives are effective to promote effective regulation; however, conditional trust may deteriorate the "social pact". Establishing a virtuous feedback between users' trust and regulators' reputation thus appears to be key to nudge developers towards creating safe AI. However, the level at which this trust emerges may depend on the specific LLM used for testing. Our results thus provide guidance for AI regulation systems, and help predict the outcome of strategic LLM agents, should they be used to aid regulation itself.

博弈论AI监管LLM行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。