arXiv:2504.14325cs.AI2025-04被引 15

用博弈论框架发现大模型智能体的偏见行为

FAIRGAME: a Framework for AI Agents Bias Recognition using Game Theory

  • 构建游戏化框架,分析多智能体互动中的策略行为
  • 揭示不同大模型、语言和性格设定导致的偏见结果
  • 适合研究智能体公平性与战略决策的学者使用

让人工智能智能体在多智能体应用中交互,增加了对AI结果可解释性和预测性的复杂性,对其在科研与社会中的可信采用具有深远影响。博弈论为捕捉和解释智能体间的策略互动提供了强大模型,但需要可复现、标准化且用户友好的信息技术框架来支持结果对比与解释。为此,我们提出FAIRGAME——一种基于博弈论的AI智能体偏见识别框架。本文描述其实现与使用方法,并利用它揭示了在主流智能体游戏中,偏见结果随所用大型语言模型(LLM)、语言、智能体人格特质或策略知识的不同而变化。总体而言,FAIRGAME使用户能够可靠且便捷地模拟所需的游戏与场景,跨模拟实验和博弈论预测进行结果比较,从而系统性发现偏见、预测策略互动中涌现的行为,并推动基于LLM智能体的战略决策研究。

原文摘要 · Abstract (English)

Letting AI agents interact in multi-agent applications adds a layer of complexity to the interpretability and prediction of AI outcomes, with profound implications for their trustworthy adoption in research and society. Game theory offers powerful models to capture and interpret strategic interaction among agents, but requires the support of reproducible, standardized and user-friendly IT frameworks to enable comparison and interpretation of results. To this end, we present FAIRGAME, a Framework for AI Agents Bias Recognition using Game Theory. We describe its implementation and usage, and we employ it to uncover biased outcomes in popular games among AI agents, depending on the employed Large Language Model (LLM) and used language, as well as on the personality trait or strategic knowledge of the agents. Overall, FAIRGAME allows users to reliably and easily simulate their desired games and scenarios and compare the results across simulation campaigns and with game-theoretic predictions, enabling the systematic discovery of biases, the anticipation of emerging behavior out of strategic interplays, and empowering further research into strategic decision-making using LLM agents.

智能体博弈论偏见检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。