智能体通过找食物间接聚集,无需直接奖励抱团。
Emergent aggregation from collective foraging

- 让智能体只感知同伴,靠找食物的个人奖励优化行为。
- 视觉范围增大时,从单体搜索转为集体搜索并出现聚集。
- 揭示资源驱动是集体行为涌现的通用机制,适合对群体智能感兴趣者。
生物系统中的集体行为通常被建模为直接社会驱动力的结果:个体因与邻居对齐或靠近而获益。本文展示,聚集也可源于间接目标。我们让强化学习觅食者初始执行随机游走,仅基于找到可再生目标的个人奖励优化行为,且只能感知同类,无法感知目标。随着视野扩大,智能体在环境适应性个体搜索与尺度无关的集体搜索之间发生急剧转变,此转变恰与空间聚集现象同步出现。因此,集体相作为最优觅食的副产品自然产生,无需直接奖励抱团。一个最小化的首次通过分析模型再现了这一转变,表现为两种搜索策略间的跨界。结果表明,间接的资源驱动奖励是集体现象涌现的普遍路径。
原文摘要 · Abstract (English)
Collective behaviour in living systems is usually modelled as the outcome of a \emph{direct} social drive: agents are rewarded, or hard-wired, to align with or approach their neighbours. Here we show that aggregation can instead emerge from an \emph{indirect} objective. We let reinforcement learning foragers, initially performing a random walk, optimize their dynamics from a purely individual reward for finding replenishable targets, while perceiving only their conspecifics and never the targets themselves. As the visual range grows, the agents undergo a sharp crossover from an environment-tuned individual search to a scale-agnostic collective one, and this crossover coincides with the onset of spatial aggregation. Thus a collective phase arises as a by-product of optimal foraging, without any direct reward for grouping. A minimal analytical first-passage model reproduces the transition as a crossover between the two search strategies. Our results identify indirect, resource-driven reward as a generic route to emergent collective phenomena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。