arXiv:2607.11913cs.NEcs.AI2026-07

受人类成长规律启发,动态调节遗传网络的探索与利用平衡。

Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming

  • 模仿人类从试错到反思的成长过程,设计自适应进化机制
  • 在Tileworld任务中显著提升智能体策略性能,最优组合超越基准37%
  • 适用于各类遗传网络变体,无需调整参数即可增强演化效率

近年来,基于图的代理智能系统因具备可解释性、以人为中心及非线性推理能力而受到关注。遗传网络编程(GNP)是一种自演化算法,通过有向图生成可解释的决策结构。如同多数演化算法,探索与利用的平衡是关键挑战,但现有研究对此关注不足。本文受人类发展规律启发——儿童偏好广泛实验与行动,随年龄增长转向深思熟虑。我们将GNP中的判断节点映射为思考,处理节点映射为行动,提出人本启发的自适应框架HGNP,通过新型自适应交叉与变异算子及循环消除机制,动态调控演化过程中的探索-利用平衡。该方法不仅提升演化效率,还能根据环境特性自动调节平衡策略。相比传统通过固定概率调参的方式,效果更优。该改进具有普适性,可应用于几乎所有GNP变体。在Tileworld基准测试中,结合标准GNP及两种最新变体均表现优异,其中与情境式GNP(HGNP-SBGNP)的组合取得最佳整体结果。

原文摘要 · Abstract (English)

Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-linear reasoning workflows. A prominent example is Genetic Network Programming (GNP), a self-evolving algorithm that utilizes directed graphs to evolve interpretable decision structures for agents. As in most evolutionary algorithms, effectively balancing exploration and exploitation is a key aspect of GNP. However, this trade-off has received limited attention in the GNP literature. To address this gap, we draw inspiration from human developmental patterns, where children prioritize broad experimentation and action over deliberation, with this tendency reversing with age. By mapping transitions between GNP's judgment nodes to deliberation and processing nodes to action, we propose Human-Inspired GNP (HGNP), a novel adaptive framework that dynamically regulates the exploration-exploitation balance throughout the evolutionary process. The method consists of novel adaptive crossover and mutation operators, and a cycle elimination mechanism. HGNP not only improves the evolutionary process but also provides a framework for adjusting the exploration-exploitation balance based on the characteristics of the target environment and its search space. This approach is more effective than tuning via crossover and mutation probabilities in standard GNP. The modifications are general and can be applied to almost all GNP variants. When integrated with standard GNP and two recently introduced GNP variants and evaluated on the Tileworld benchmark, HGNP demonstrated significant performance improvement in agents' strategy. The combination of HGNP with Situation-based GNP (HGNP-SBGNP) achieved the best overall results.

遗传网络自演化探索利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。