arXiv:2601.04365cs.LG2026-01

程序化策略比神经网络策略在演化强化学习中存活更久。

Survival Dynamics of Neural and Programmatic Policies in Evolutionary Reinforcement Learning

  • 用可微分决策列表实现程序化策略,提升行为可解释性。
  • 程序化策略平均多存活201.69步,显著优于神经网络策略。
  • 仅靠学习的程序化策略比兼具学习与评估的神经策略多活73.67步。

在演化强化学习(ERL)任务中,智能体策略常以小型人工神经网络(NERL)编码,缺乏显式模块结构,限制了行为可解释性。本文研究程序化策略(PERL)是否能媲美NERL性能。为此,我们首次完整复现并开源了经典的1992年人工生命(ALife)ERL测试环境。通过4000次独立试验,采用Kaplan-Meier曲线和受限均生存时间(RMST)进行严谨生存分析,结果表明PERL与NERL在生存概率上存在统计显著差异:PERL智能体平均比NERL多存活201.69步。此外,仅使用学习机制的SDDL(软可微分决策列表)智能体,平均比同时使用学习与评估的神经智能体多存活73.67步。这些结果证明,在人工生命场景下,程序化策略可超越神经策略的生存表现。

原文摘要 · Abstract (English)

In evolutionary reinforcement learning tasks (ERL), agent policies are often encoded as small artificial neural networks (NERL). Such representations lack explicit modular structure, limiting behavioral interpretation. We investigate whether programmatic policies (PERL), implemented as soft, differentiable decision lists (SDDL), can match the performance of NERL. To support reproducible evaluation, we provide the first fully specified and open-source reimplementation of the classic 1992 Artificial Life (ALife) ERL testbed. We conduct a rigorous survival analysis across 4000 independent trials utilizing Kaplan-Meier curves and Restricted Mean Survival Time (RMST) metrics absent in the original study. We find a statistically significant difference in survival probability between PERL and NERL. PERL agents survive on average 201.69 steps longer than NERL agents. Moreover, SDDL agents using learning alone (no evolution) survive on average 73.67 steps longer than neural agents using both learning and evaluation. These results demonstrate that programmatic policies can exceed the survival performance of neural policies in ALife.

演化强化学习程序化策略生存分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。