用进化视角重构强化学习的三大信条,让智能体更像真实生命。
Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light
- 以开放式新颖性搜索替代传统优化,模拟生物适应机制
- 引入热力学理论解释生命起源与自主性的关系
- 适合对具身智能、生命本质感兴趣的学者
人工学习系统正从被动学习者转向日益自主的智能体,这使得‘什么是代理’的问题变得尤为紧迫。强化学习(RL)提供了最明确的智能体-环境互动框架,基于三大核心信条:环境是马尔可夫决策过程、学习即策略优化、智能体是标量奖励的最大化者。近期研究呼吁重新审视这些信条:将学习理解为适应而非优化,拓展目标超越标量奖励,指出在高度强调环境的形式体系中缺乏对智能体的正式理论。我们主张,人工生命领域有独特优势来照亮这一批判,并构建替代方案。通过借鉴开放式新颖性搜索作为适应与目标导向行为的互补模型,将此类进化动态建立在生命起源与代理性的热力学理论基础上,从而提出一种更符合生物学且形式上严谨的自适应智能体定义。
原文摘要 · Abstract (English)
Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the question of what constitutes agency. Reinforcement learning (RL) offers arguably the most explicit formulation of agent-environment interaction, built on three core tenets: the environment as a Markov decision process, learning as policy optimization, and the agent as a maximizer of scalar reward. Recent work has called to revise these tenets: reconceptualizing learning as adaptation rather than optimization, broadening goals beyond scalar reward, and noting the absence of a formal theory of the agent in a formalism that so heavily emphasizes the environment. We argue that the artificial life community is uniquely positioned to illuminate this critique and concretize an alternative. We draw on open-ended novelty search as a complementary model of adaptation and goal-directed behavior beyond reward optimization, and ground such evolutionary dynamics in thermodynamic theories of origin-of-life and agency, toward a more biologically faithful and formally grounded account of what it is to be an adaptive agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。