arXiv:2604.03582cs.LG2026-04被引 3

用低秩注意力提升神经算子对物理方程的建模效率

Simple yet Effective: Low-Rank Spatial Attention for Neural Operators

论文配图:Simple yet Effective: Low-Rank Spatial Attention for Neural Operators
图 1 · 摘自论文原文
  • 通过低秩压缩实现全局空间交互的高效建模
  • 平均误差降低17%以上,优于现有最优方法
  • 仅用标准Transformer组件,易实现且适合硬件加速

神经算子作为求解偏微分方程(PDE)的数据驱动代理已崭露头角,其成功关键在于高效建模由物理规律引发的空间点间长程全局耦合。在多数PDE场景中,诱导出的全局相互作用核具有可压缩性,呈现快速谱衰减,支持低秩近似。本文基于此观察,统一现有神经算子中的全局混合模块,提出一个共享的低秩模板:将高维逐点特征压缩至紧凑潜在空间,在其中处理全局交互,再重建回空间点。据此提出低秩空间注意力(LRSA),以标准Transformer组件(注意力、归一化、前馈网络)构建,结构简洁,易于实现且兼容硬件优化内核。实验表明,该简单设计即可达到高精度,平均误差相对最优方法降低超过17%,同时在混合精度训练下保持稳定高效。

原文摘要 · Abstract (English)

Neural operators have emerged as data-driven surrogates for solving partial differential equations (PDEs), and their success hinges on efficiently modeling the long-range, global coupling among spatial points induced by the underlying physics. In many PDE regimes, the induced global interaction kernels are empirically compressible, exhibiting rapid spectral decay that admits low-rank approximations. We leverage this observation to unify representative global mixing modules in neural operators under a shared low-rank template: compressing high-dimensional pointwise features into a compact latent space, processing global interactions within it, and reconstructing the global context back to spatial points. Guided by this view, we introduce Low-Rank Spatial Attention (LRSA) as a clean and direct instantiation of this template. Crucially, unlike prior approaches that often rely on non-standard aggregation or normalization modules, LRSA is built purely from standard Transformer primitives, i.e., attention, normalization, and feed-forward networks, yielding a concise block that is straightforward to implement and directly compatible with hardware-optimized kernels. In our experiments, such a simple construction is sufficient to achieve high accuracy, yielding an average error reduction of over 17\% relative to second-best methods, while remaining stable and efficient in mixed-precision training.

神经算子注意力机制低秩近似PDE求解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。