arXiv:2410.17517cs.MAcs.AI2024-10被引 1

群体模仿行为可等效为单一强化学习代理,解释集体智能形成机制。

The Hive Mind is a Single Reinforcement Learning Agent

  • 基于蜜蜂摇摆舞的模仿投票模型,等效于多臂赌博机的在线强化学习。
  • 群体学习更新规则等同于新提出的‘梅纳德-克罗斯学习’算法。
  • 适用于研究生物、经济与社会系统中的集体决策,启发人工协同学习设计。

决策是智能体或群体的核心属性。自然界通过两种不同机制实现有效策略:个体间通过模仿进行集体决策,或由单个智能体通过试错学习。本文建立这两种范式的等价关系:遵循简单局部模仿规则的个体群体所涌现的分布式认知(常称为‘蜂群心智’),等效于一个聚合多个并行多臂赌博机动作价值样本的单一在线强化学习(RL)代理。具体而言,我们证明蜜蜂摇摆舞的盲目模仿加权投票模型中,宏观代理的学习更新规则是一种新型强化学习算法——‘梅纳德-克罗斯学习’。分析表明,遵循简单模仿策略的个体群体可等效于更复杂的智能学习实体,支持了群体智能可解释自然中看似非理性个体行为选择的观点。该框架在生物学之外还可用于分析经济与社会系统中个体对成功策略的模仿,构成一种集体学习过程。研究结果亦可指导人工领域中受强化学习启发的可扩展集体系统设计。

原文摘要 · Abstract (English)

Decision-making is an essential attribute of any intelligent agent or group. Natural systems are known to converge to effective strategies through at least two distinct mechanisms: collective decision-making via imitation of others, and trial-and-error by a single agent. This paper establishes an equivalence between these two paradigms. We show that the emergent distributed cognition (sometimes referred to as the \textit{hive mind}) arising from individuals following simple, local imitation-based rules is that of a single online reinforcement learning (RL) agent aggregating action-value samples from many parallel instances of a multi-armed bandit. More specifically, we show that, in the blindly imitative weighted voter model of honey bees' waggle dance, the update rule through which this macro-agent learns is an RL algorithm that we coin \textit{Maynard-Cross Learning}. Our analysis implies that a group of individuals following simple imitative strategies can be equivalent to a more complex learning entity, substantiating the idea that group-level intelligence may explain how seemingly irrational individual behaviors are selected in nature. Beyond biology, the framework offers new tools for analyzing economic and social systems where individuals imitate successful strategies, effectively participating in a collective learning process. Our findings may further inform the design of scalable RL-inspired collective systems in artificial domains.

强化学习群体智能集体决策生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。