arXiv:2410.05254cs.CLcs.AI2024-10被引 25

构建语言经济博弈统一框架,评估大模型在谈判中的理性与公平表现。

GLEE: A Unified Framework and Benchmark for Language-based Economic Environments

  • 设计三类可参数化的双人语言博弈,标准化经济行为研究。
  • 发现市场参数与模型选择对结果有复杂耦合影响,非线性交互显著。
  • 开源框架支持大模型间及人机交互实验,适合智能系统伦理与机制设计研究者。

大型语言模型(LLMs)在以自然语言沟通为主的经济与战略互动中展现出巨大潜力。这引发关键问题:LLMs是否表现出理性?其表现相较于人类如何?能否达成高效且公平的结果?自然语言在策略互动中扮演什么角色?经济环境特征如何影响这些动态?这些问题对将基于LLM的智能体集成到在线零售平台、推荐系统等现实数据驱动系统具有重要经济与社会意义。为此,我们提出一个标准化的双人、顺序、语言驱动博弈基准。受经济学文献启发,定义了三类具有统一参数化、自由度和经济度量(如自我收益)的游戏家族,用于评估代理行为及博弈结果(效率与公平性)。我们开发了一个开源交互模拟与分析框架,并收集了大量LLM vs. LLM的交互数据,以及额外的人类 vs. LLM交互数据集。通过广泛实验,验证该框架可实现:(i) 在不同经济情境下比较LLM代理的行为;(ii) 评估代理在个体与集体绩效上的表现;(iii) 量化环境经济特征对代理行为的影响。结果显示,市场参数与模型选择对经济结果具有复杂且相互依赖的影响,凸显需谨慎设计与分析语言驱动的经济生态系统。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show significant potential in economic and strategic interactions, where communication via natural language is often prevalent. This raises key questions: Do LLMs behave rationally? How do they perform compared to humans? Do they tend to reach an efficient and fair outcome? What is the role of natural language in strategic interaction? How do characteristics of the economic environment influence these dynamics? These questions become crucial concerning the economic and societal implications of integrating LLM-based agents into real-world data-driven systems, such as online retail platforms and recommender systems. To answer these questions, we introduce a benchmark for standardizing research on two-player, sequential, language-based games. Inspired by the economic literature, we define three base families of games with consistent parameterization, degrees of freedom and economic measures to evaluate agents' performance (self-gain), as well as the game outcome (efficiency and fairness). We develop an open-source framework for interaction simulation and analysis, and utilize it to collect a dataset of LLM vs. LLM interactions across numerous game configurations and an additional dataset of human vs. LLM interactions. Through extensive experimentation, we demonstrate how our framework and dataset can be used to: (i) compare the behavior of LLM-based agents in various economic contexts; (ii) evaluate agents in both individual and collective performance measures; and (iii) quantify the effect of the economic characteristics of the environments on the behavior of agents. Our results suggest that the market parameters, as well as the choice of the LLMs, tend to have complex and interdependent effects on the economic outcome, which calls for careful design and analysis of the language-based economic ecosystem.

语言模型经济博弈人机交互基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。