arXiv:2603.12276cs.LG2026-03

用几何启发的核函数替代传统神经网络结构,提升稳定性与泛化能力。

No More DeLuLu: Physics-Inspired Kernel Networks for Geometrically-Grounded Neural Computation

  • 提出 yat-product 核运算,结合二次对齐与反平方距离机制。
  • 在 MNIST 上分类性能媲美线性模型,且原型演化有界、抗干扰强。
  • 适用于图像与语言任务,适合追求稳定性和可解释性的研究者。

我们提出 yat-product,一种结合二次对齐与反平方接近性的核算子。证明其为 Mercer 核、解析函数、在有界域上 Lipschitz 连续,并具有自正则性,可唯一嵌入再生核希尔伯特空间(RKHS)。神经物质网络(NMNs)仅使用 yat-product 作为非线性单元,取代传统线性-激活-归一化模块,将归一化功能内置于核的分母中,无需独立归一化层。该架构简化后仍保持通用逼近能力。实验表明,基于 NMN 的分类器在 MNIST 上性能与线性基线相当,原型演化有界且具备超位置鲁棒性;在语言建模中,Aether-GPT2 使用相同参数量时验证损失低于 GPT-2,采用 yat 基注意力与 MLP 模块。本框架统一了核学习、梯度稳定性与信息几何,确立了 NMN 作为传统神经架构的原理性替代方案。

原文摘要 · Abstract (English)

We introduce the yat-product, a kernel operator combining quadratic alignment with inverse-square proximity. We prove it is a Mercer kernel, analytic, Lipschitz on bounded domains, and self-regularizing, admitting a unique RKHS embedding. Neural Matter Networks (NMNs) use yat-product as the sole non-linearity, replacing conventional linear-activation-normalization blocks with a single geometrically-grounded operation. This architectural simplification preserves universal approximation while shifting normalization into the kernel itself via the denominator, rather than relying on separate normalization layers. Empirically, NMN-based classifiers match linear baselines on MNIST while exhibiting bounded prototype evolution and superposition robustness. In language modeling, Aether-GPT2 achieves lower validation loss than GPT-2 with a comparable parameter budget while using yat-based attention and MLP blocks. Our framework unifies kernel learning, gradient stability, and information geometry, establishing NMNs as a principled alternative to conventional neural architectures.

神经网络核方法几何计算GPT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。