对比真实老鼠与强化学习智能体在避险任务中的表现,发现智能体缺乏自我保护本能。
Of Mice and Machines: A Comparison of Learning Between Real World Mice and RL Agents
- 设计避险迷宫环境,对比生物老鼠与强化学习智能体的行为差异
- 智能体为微小效率提升甘愿冒死,而老鼠会谨慎评估风险并规避危险
- 提出两种新机制,使智能体行为更接近真实生物的避险策略
强化学习(RL)在复杂决策任务中取得显著进展,引发一个自然问题:这些人工系统与经过数百万年进化塑造的生物体相比如何?为回答此问题,我们开展了一项比较研究,考察生物老鼠与强化学习智能体在捕食者回避迷宫环境中的表现。分析发现显著差异:强化学习智能体始终缺乏自我保护本能,为微小效率提升而轻易冒险‘死亡’;而生物体则展现出复杂的风险评估与规避行为。为弥合这一差距,我们提出了两种促进更自然风险规避行为的新机制。该方法促使智能体涌现出自然行为模式,包括战略性环境评估、谨慎路径规划以及与生物体高度相似的捕食者回避策略。
原文摘要 · Abstract (English)
Recent advances in reinforcement learning (RL) have demonstrated impressive capabilities in complex decision-making tasks. This progress raises a natural question: how do these artificial systems compare to biological agents, which have been shaped by millions of years of evolution? To help answer this question, we undertake a comparative study of biological mice and RL agents in a predator-avoidance maze environment. Through this analysis, we identify a striking disparity: RL agents consistently demonstrate a lack of self-preservation instinct, readily risking ``death'' for marginal efficiency gains. These risk-taking strategies are in contrast to biological agents, which exhibit sophisticated risk-assessment and avoidance behaviors. Towards bridging this gap between the biological and artificial, we propose two novel mechanisms that encourage more naturalistic risk-avoidance behaviors in RL agents. Our approach leads to the emergence of naturalistic behaviors, including strategic environment assessment, cautious path planning, and predator avoidance patterns that closely mirror those observed in biological systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。