神经进化棋手在学习中演化出个性,因想象放大差异。
Personality Requires Struggle: Three Regimes of the Baldwin Effect in Neuroevolved Chess Agents
- 用神经进化+海布学习,在棋局中模拟终身学习
- 34代后行为多样性反超,62%走法不一致
- 适合研究人格演化与自我对弈系统缺陷
寿命学习能否在进化时间尺度上扩展行为多样性而非压缩它?我们测试了竞争性领域中的神经进化棋手,其包含八个NEAT演化的神经模块、局内海布学习机制以及带有想象力的可取性信号链。每种海布条件运行10个种子,结果发现方差交叉现象:海布开启初期跨种子方差低于关闭状态,但在第34代后反超。该趋势单调(ρ=0.91,p<10⁻⁶),表明可塑性对行为方差的影响随进化逆转——初期压缩多样性(符合旧理论),后期则通过想象力放大进化出的感知差异,形成反馈循环,这是突变无法维持的。最终实现结构化行为分化:同一局面选择不同着法(62%分歧)、发展出不同开局套路、子力偏好和对局时长。这些不是随机采样策略,而是可复现的行为特征(ICC>0.8)且信号链配置可解释。根据对手类型呈现三种模式:探索型(海布开,对手多样)、彩票型(海布关,精英锁定)、透明型(同模型对手,大脑自消)。透明型提出可验证预测:自对弈系统可能因消除异质性而系统抑制行为多样性,而这正是人格所需。
原文摘要 · Abstract (English)
Can lifetime learning expand behavioral diversity over evolutionary time, rather than collapsing it? Prior theory predicts that plasticity reduces variance by buffering organisms against environmental noise. We test this in a competitive domain: chess agents with eight NEAT-evolved neural modules, Hebbian within-game plasticity, and a desirability-domain signal chain with imagination. Across 10~seeds per Hebbian condition, a variance crossover emerges: Hebbian ON starts with lower cross-seed variance than OFF, then surpasses it at generation~34. The crossover trend is monotonic (\r{ho} = 0.91, p < 10^{-6): plasticity's effect on behavioral variance reverses over evolutionary time, initially compressing diversity (consistent with prior predictions) then expanding it as evolved Perception differences are amplified through imagination -- a feedback loop that mutation alone cannot sustain. The result is structured behavioral divergence: evolved agents select different moves on the same positions (62\% disagreement), develop distinct opening repertoires, piece preferences, and game lengths. These are not different sampling policies -- they are reproducible behavioral signatures (ICC > 0.8) with interpretable signal chain configurations. Three regimes appear depending on opponent type: exploration (Hebbian ON, heterogeneous opponent), lottery (Hebbian OFF, elitism lock-in), and transparent (same-model opponent, brain self-erasure). The transparent regime generates a falsifiable prediction: self-play systems may systematically suppress behavioral diversity by eliminating the heterogeneity that personality requires. \textbf{Keywords: Baldwin Effect, neuroevolution, NEAT, Hebbian learning, chess, cognitive architecture, personality emergence, imagination
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。