arXiv:2603.13263cs.LG2026-03

用球面核算子替代注意力,让世界模型更自适应且抗维度诅咒。

Beyond Attention: True Adaptive World Models via Spherical Kernel Operator

  • 引入球面核算子(SKO),通过超球面投影与广义勒让德多项式重构目标函数
  • 在语言建模中收敛速度更快,预测误差随流形维数增长而非环境维数
  • 数学上解耦环境动态与观测频率偏差,适合高维复杂系统建模

基于世界模型的人工智能普遍将高维观测投影到参数化隐空间,再学习转移动态。但该范式存在数学缺陷:仅将流形学习问题转移到隐空间。当数据分布变化时,隐流形随之偏移,迫使预测算子重新学习新拓扑结构。此外,根据经典逼近理论,如点积注意力等正算子不可避免地出现饱和现象,永久限制其预测能力,且易受维度诅咒影响。本文提出一种数学严谨的世界模型构建范式,重新定义核心预测机制。受Ryan O'Dowd工作启发,提出球面核算子(SKO),以标准注意力的替代方案。通过将未知数据流形投影至统一的超球面,并利用局部化的超球多项式(Gegenbauer),SKO实现对目标函数的直接积分重构。由于该局部球面多项式核非严格正,可规避饱和现象,逼近误差仅依赖于内在流形维数q,而非环境维数。同时,将未归一化输出形式化为真实测度支撑估计器,数学上解耦了真实的环境转移动态与代理观测频率的偏差。实证评估表明,SKO在自回归语言建模中显著加速收敛,并优于标准注意力基线。

原文摘要 · Abstract (English)

The pursuit of world model based artificial intelligence has predominantly relied on projecting high-dimensional observations into parameterized latent spaces, wherein transition dynamics are subsequently learned. However, this conventional paradigm is mathematically flawed: it merely displaces the manifold learning problem into the latent space. When the underlying data distribution shifts, the latent manifold shifts accordingly, forcing the predictive operator to implicitly relearn the new topological structure. Furthermore, by classical approximation theory, positive operators like dot product attention inevitably suffer from the saturation phenomenon, permanently bottlenecking their predictive capacity and leaving them vulnerable to the curse of dimensionality. In this paper, we formulate a mathematically rigorous paradigm for world model construction by redefining the core predictive mechanism. Inspired by Ryan O'Dowd's foundational work we introduce Spherical Kernel Operator (SKO), a framework that replaces standard attention. By projecting the unknown data manifold onto a unified ambient hypersphere and utilizing a localized sequence of ultraspherical (Gegenbauer) polynomials, SKO performs direct integral reconstruction of the target function. Because this localized spherical polynomial kernel is not strictly positive, it bypasses the saturation phenomenon, yielding approximation error bounds that depend strictly on the intrinsic manifold dimension q, rather than the ambient dimension. Furthermore, by formalizing its unnormalized output as an authentic measure support estimator, SKO mathematically decouples the true environmental transition dynamics from the biased observation frequency of the agent. Empirical evaluations confirm that SKO significantly accelerates convergence and outperforms standard attention baselines in autoregressive language modeling.

世界模型球面核注意力替代流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。