揭示神经网络集成理论在开放系统中的缺失,提出其与核反应理论的对应关系。
Integrating Out, Twice:The Open-System Case That Neural-Network Ensemble Theory Is Missing
- 用高斯代数和分块求逆,统一闭合与开放系统的数学结构
- 实测三种模型后发现开放态的物理守恒量无法有效提取,结果多为负
- 强调开放系统需连续谱与波动力学,现有学习模型不满足此条件
将神经网络参数平均与高斯子块消去视为同一操作——即施尔补运算;闭合情形下生成协方差及其逆。但开放情形缺失,而核反应理论已解决。将散射问题投影至选定通道,其余概率不可逆地流向连续谱,留下非厄米有效生成元,精确记录损失内容:即核光学模型与广义光学定理。本文仅用分布矩、高斯代数与分块求逆,对比两情形,完整给出闭合情形的映射关系:神经正切核即费舍尔敏感性核,无限宽高斯极限即高斯过程模拟器,懒惰到特征的过渡即降基模拟器的有效边界。随后在截断注意力图、标记级转移算子与稀疏专家路由器上测试开放输出,结果普遍负面。守恒通量账本虽可迁移至真正开放处,但其独特内容或为分区选择的人为产物,或被训练目标压制于下限附近。实际有用的不确定性实为认识论型,存在于对应关系的闭合半边,而非开放侧。该负结果具有结构性根源:开放情形需具备连续谱与波动动力学的消除子系统,主流学习中的有限或耗散对象无法提供。本文为注记而非成果,核心发现即此负结果,价值在于精准定位该问题所在。
原文摘要 · Abstract (English)
Averaging a neural network over its random parameters and marginalizing a Gaussian sector are the same operation, the Schur complement of the eliminated block, and when that block is closed it returns a covariance and its inverse. That is all a network ensemble produces, the closed case. The open case is missing, and nuclear reaction theory has it worked out. Projecting a scattering problem onto a chosen set of channels, with the rest carrying probability irreversibly to a continuum, leaves a non-Hermitian effective generator that conserves and itemizes exactly what it loses: the nuclear optical model and its generalized optical theorem. I set the two cases side by side using only the moments of a distribution, the algebra of Gaussians, and block inversion, no field theory, and give the closed-case dictionary in full: the neural tangent kernel is the Fisher sensitivity kernel, the infinite-width Gaussian limit is the Gaussian-process emulator, and the lazy-to-feature transition is the validity boundary of a reduced-basis emulator. I then test the open export on a truncated attention map, a token-level transfer operator, and a sparse expert router, and report a mostly negative result. The conserved flux ledger ports wherever openness is genuinely present, but its distinctive content is absent, an artifact of the chosen partition, or pinned near a floor by the training objective, and the operationally useful uncertainty turns out to be epistemic, living in the closed half of the correspondence, not the open one. The negative has a structural reason this note makes precise: the open case needs an eliminated sector with a continuous spectrum and wave-like, not relaxational, dynamics, which mainstream learning's finite or dissipative objects do not supply. This is a note, not a result; its main finding is that negative one, and its value is the map that locates it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。