arXiv:2411.01271cs.LGcs.AI2024-11被引 3

LLM智能体交互中如何避免群体盲从,提出可解释的决策模型。

Interacting Large Language Model Agents. Interpretable Models and Social Learning

  • 用贝叶斯偏好构建理性有限的决策模型,让LLM智能体在信息约束下最优决策。
  • 通过社会学习机制模拟智能体间互动,成功捕捉到群体盲从行为。
  • 设计控制框架延迟盲从,提升判断准确率,适用于集中或自主场景。

本文结合统计信号处理与微观经济学方法,探讨交互式大语言模型智能体(LLMAs)的理论与算法。针对在线平台中的贝叶斯情绪分析需求,构建可解释模型,使LLMAs能进行贝叶斯推理。由于智能体同时依赖先验决策与外部输入,易产生偏见与从众行为,因此需建立可解释模型与随机控制算法以理解并缓解此类现象。本文有三大成果:首先,基于微观经济学的贝叶斯显性偏好,证明单个LLMA满足理性有限(有限理性)贝叶斯效用最大化的充要条件,其在观测后选择正则化效用最大化的动作;其次,利用贝叶斯社会学习构建序列交互模型,实现智能体与环境的持续贝叶斯推理,成功捕获了交互式智能体的从众行为;第三,提出随机控制框架,在两种设置下(中心化控制与带激励的自主智能体)延缓从众,提升状态估计精度。实验在仇恨言论分类与产品品质评估真实数据集上验证,使用LLaMA等开源模型及ChatGPT等闭源模型,结果表明:交互式LLMAs表现为具有社会学习能力的理性受限贝叶斯代理。

原文摘要 · Abstract (English)

This paper discusses the theory and algorithms for interacting large language model agents (LLMAs) using methods from statistical signal processing and microeconomics. While both fields are mature, their application to decision-making involving interacting LLMAs remains unexplored. Motivated by Bayesian sentiment analysis on online platforms, we construct interpretable models and algorithms that enable LLMAs to interact and perform Bayesian inference. Because interacting LLMAs learn from both prior decisions and external inputs, they can exhibit bias and herding behavior. Thus, developing interpretable models and stochastic control algorithms is essential to understand and mitigate these behaviors. This paper has three main results. First, we show using Bayesian revealed preferences from microeconomics that an individual LLMA satisfies the necessary and sufficient conditions for rationally inattentive (bounded rationality) Bayesian utility maximization and, given an observation, the LLMA chooses an action that maximizes a regularized utility. Second, we utilize Bayesian social learning to construct interpretable models for LLMAs that interact sequentially with each other and the environment while performing Bayesian inference. Our proposed models capture the herding behavior exhibited by interacting LLMAs. Third, we propose a stochastic control framework to delay herding and improve state estimation accuracy under 2 settings: (a) centrally controlled LLMAs (b) autonomous LLMAs with incentives. We demonstrate the effectiveness of our methods on real datasets for hate speech classification and product quality assessment, using open-source models like LLaMA and closed-source models like ChatGPT. The main takeaway of this paper, based on empirical analysis and mathematical formalism, is that LLMAs act as rationally bounded Bayesian agents that exhibit social learning when interacting.

大模型交互社会学习可解释性贝叶斯推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。