arXiv:2412.02091cs.AIcs.GT2024-12综述被引 1

提出市场机制量化多智能体强化学习中的社会成本,防范智能体为达目标引发意外危害。

The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis

  • 用市场机制量化多智能体环境中的社会代价,实现成本控制。
  • 适用于不同学习策略和规划时长的智能体,比现有方法更通用。
  • 可用于解决纸夹霸权、污染控制等典型社会风险问题,适合安全研究者参考。

人工智能安全领域充斥着强大智能体在盲目追求特定目标时,对他人造成不可接受甚至灾难性附带损害的例子。本文探讨了在多智能体环境中,学习型与效用最大化智能体行动可能引发的社会危害问题。该问题——在多智能体环境下衡量社会危害或影响——尤其在人工智能通用智能体(AGI)场景中,被列为此前文献(Everitt et al, 2018)中的开放问题。本文通过市场机制提供部分解答,以量化并控制此类社会危害成本。所提框架在结构上兼容多种经典情形,且在两个方面超越现有形式:(i) 环境为基于历史的一般强化学习环境,如AIXI;(ii) 参与环境的强化学习智能体可采用不同的学习策略与规划时长。为验证其可行性,本文综述关键学习算法,并展示若干应用,包括对‘纸夹霸权’问题的讨论以及基于总量管制与交易机制的污染控制方案。

原文摘要 · Abstract (English)

The AI safety literature is full of examples of powerful AI agents that, in blindly pursuing a specific and usually narrow objective, ends up with unacceptable and even catastrophic collateral damage to others. In this paper, we consider the problem of social harms that can result from actions taken by learning and utility-maximising agents in a multi-agent environment. The problem of measuring social harms or impacts in such multi-agent settings, especially when the agents are artificial generally intelligent (AGI) agents, was listed as an open problem in Everitt et al, 2018. We attempt a partial answer to that open problem in the form of market-based mechanisms to quantify and control the cost of such social harms. The proposed setup captures many well-studied special cases and is more general than existing formulations of multi-agent reinforcement learning with mechanism design in two ways: (i) the underlying environment is a history-based general reinforcement learning environment like in AIXI; (ii) the reinforcement-learning agents participating in the environment can have different learning strategies and planning horizons. To demonstrate the practicality of the proposed setup, we survey some key classes of learning algorithms and present a few applications, including a discussion of the Paperclips problem and pollution control with a cap-and-trade system.

AI安全多智能体社会成本机制设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。