arXiv:2609.04407cs.LG2026-09

通过可控实验揭示注意力机制如何提升神经算子的PDE求解精度。

Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures

论文配图:Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures
图 1 · 摘自论文原文
  • 对比五种注意力结构,分离分析各自对精度的影响。
  • 跨注意力+传感器分块使误差降低2.4至28倍,最优配置达3.5至32.3倍。
  • 查询依赖的跨注意力最可靠,分支自注意力适合复杂空间输入。

深度神经算子可学习输入函数到完整偏微分方程(PDE)解场的映射,实现新问题实例的前向求解速度比传统数值求解器快数个数量级。近期研究将注意力机制引入神经算子,但多数工作同时改变多个架构组件,难以识别准确率提升的真实原因。本文系统比较五种具不同注意力机制的DeepONet变体,在数据驱动与物理信息双重训练范式下,分离分析跨注意力、自注意力、标记化及注意力深度的影响。评估任务包括:源驱动的一维非线性扩散-反应方程、初始条件可变的一维黏性Burgers方程,以及具有异质源场的二维泊松热传导问题。基于传感器的分块标记结合跨注意力,使经典DeepONet在所有基准训练组合中均降低2.4–28.0倍的平均相对L₂误差;最佳配置下误差降低3.5–32.3倍。仅使用点积融合的分支自注意力表现不稳定,在一维问题中恶化性能,但在更复杂的二维源场问题中有帮助;若叠加于跨注意力之上,则在六组测试中均有提升,但增益小于仅用跨注意力融合。全局预混合未带来一致优势。增加跨注意力深度可进一步提升精度,但收益递减且在物理信息训练中成本显著上升。总体而言,查询依赖的跨注意力最为可靠,而分支自注意力最适合大尺度、空间复杂的函数输入。

原文摘要 · Abstract (English)

Deep neural operators learn mappings between input functions and complete PDE solution fields, enabling forward evaluations of new problem instances orders of magnitude faster than conventional numerical solvers. Attention mechanisms have recently been introduced into neural operators, but most studies change several architectural components at once, making it difficult to identify what actually improves accuracy. This work presents a controlled and systematic study of five deep operator network (DeepONet) variants with distinct attention mechanisms, trained under both data-driven and physics-informed regimes, to isolate the effects of cross-attention, self-attention, tokenization, and attention depth. We evaluate them on a source-driven transient one-dimensional nonlinear diffusion-reaction equation, a transient one-dimensional viscous Burgers equation with variable initial conditions, and a two-dimensional Poisson heat-conduction problem with heterogeneous source fields. Per-sensor tokenization with cross-attention reduces the mean relative L_2 error of the classical DeepONet in all benchmark-training combinations by factors of 2.4-28.0, while the best attention configurations reach 3.5-32.3. Branch self-attention paired only with dot-product fusion is inconsistent, degrading the one-dimensional problems while helping the more complex two-dimensional source field; added on top of cross-attention it improves all six cases, though by less than cross-attention fusion alone. Global pre-mixing provides no consistent benefit. Increasing cross-attention depth further improves accuracy, but with diminishing returns and a substantially higher cost under physics-informed training. Overall, query-dependent cross-attention is the most reliable mechanism, whereas branch self-attention is most useful for large, spatially complex functional inputs.

神经算子注意力机制PDE求解深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。