深度神经算子会逐渐遗忘几何结构,影响精度与泛化能力。
Do Neural Operators Forget Geometry? The Forgetting Hypothesis in Deep Operator Learning

- 通过层级几何探测发现,算子网络随深度增加逐渐丢失几何信息。
- 几何遗忘导致模型精度下降、稳定性变差、泛化能力减弱。
- 引入轻量几何记忆机制可有效缓解遗忘,尤其适用于注意力模型。
神经算子在规则域上表现良好,但在不规则几何上的行为仍不明确。我们发现,这种局限性并非编码问题,而是深层算子架构固有的深度失效模式。我们提出几何遗忘假说:由于算子层的马尔可夫结构及其对全局混合机制的依赖,随着深度增加,神经算子逐步丧失对域几何的访问能力。通过层级几何探测,我们证明了谱方法和基于注意力的算子均系统性地丢失几何保真度。我们表明,这种几何遗忘会降低准确率、稳定性和泛化能力。为应对该问题,我们引入一种轻量级几何记忆注入机制,在中间层恢复几何约束,仅带来极小的架构开销。该简单干预能持续缓解遗忘,并暴露了基于Transformer的算子中存在几何捷径不稳定性,揭示几何保留是结构性需求而非设计选择。
原文摘要 · Abstract (English)
Neural operators perform well on structured domains, yet their behaviour on irregular geometries remains poorly understood. We show that this limitation is not merely an encoding issue, but a depth-wise failure mode inherent to deep operator architectures. We formalise the Geometric Forgetting Hypothesis: due to the Markovian structure of operator layers and their reliance on global mixing mechanisms, neural operators progressively lose access to domain geometry as depth increases. Using layer-wise geometric probing, we demonstrate that both spectral and attention-based operators systematically lose geometric fidelity. We show that this geometric forgetting degrades accuracy, stability, and generalisation. To counteract it, we introduce a lightweight geometry memory injection mechanism that restores geometric constraints at intermediate depths with minimal architectural overhead. This simple intervention consistently mitigates forgetting and exposes a geometric shortcut instability in transformer-based operators, revealing that geometric retention is a structural requirement rather than a design choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。