arXiv:2608.14894cs.LG2026-08被引 1

神经网络通过自我干预实验,学会预测自身结构变化的后果。

Can Neural Networks Learn by Experimenting on Themselves? Self-Interventional Learning from Functional Consequences to Predictive Self-Knowledge

  • 让神经网络主动修改自身结构并观察结果,学习可预测的自模型。
  • 干预次数从4增至56,预测误差降低55.8%,相关性提升至0.883。
  • 适合对模型可解释性、自适应系统感兴趣的科研人员。

机器学习系统通常只建模外部数据,其内部功能结构由外部观察者分析。本文提出自干预学习(SIL),即神经网络主动扰动自身功能结构,观察结果,学习预测性自模型,泛化到未执行的干预,并用预测指导后续结构操作。在已知结构的合成系统中,SIL成功恢复了关键结构、冗余性和可替换性,但协同效应未能可靠恢复。在30个新种子测试中,将成对干预预算从4增加到56,保留预测误差从0.0335降至0.0148,Spearman相关性从0.629升至0.883。消融实验表明,保留正确的干预-后果映射使未来预测误差降低81.3%;使用相同自模型指导行为使归一化遗憾降低31.7%,相比忽略它。然而,模型引导行为并未显著优于直接经验记忆策略,且在等预算的CIFAR-10/ResNet验证中未表现出稳健性优势。结果支持SIL作为基于干预的学习框架,用于获取对网络自身功能组织的预测知识,同时表明自模型仍不完整,且不总优于简单直接策略。

原文摘要 · Abstract (English)

Machine-learning systems usually model external data, while their internal functional organization is analyzed by external observers. This work introduces Self-Interventional Learning (SIL), in which a neural system perturbs its own functional structure, observes consequences, learns a predictive self-model, generalizes to unexecuted interventions, and uses predictions to guide later structural action. In a construction-known synthetic system, SIL recovered critical structure, redundancy, and replaceability, while synergy was not reliably recovered. Across 30 fresh confirmatory seeds, increasing the pairwise intervention budget from 4 to 56 reduced held-out prediction error from 0.0335 to 0.0148 and increased Spearman correlation from 0.629 to 0.883. In a matched ablation, preserving the correct intervention--consequence mapping reduced prospective prediction error by 81.3%, while using the same learned self-model for action reduced normalized regret by 31.7% relative to ignoring it. However, model-guided action did not significantly outperform a direct empirical-memory policy, and powered CIFAR-10/ResNet validation showed no robustness advantage over equal-budget direct repair search. These results support SIL as an intervention-driven framework for learning predictive knowledge about a network's own functional organization, while showing that the self-model remains incomplete and is not universally superior to simpler direct strategies.

自学习神经网络可解释性干预学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。