arXiv:2603.22525cs.LGcs.CR2026-03被引 4

神经算子模型在核能系统中易受极稀疏干扰攻击,导致预测崩溃却无法被常规检测发现。

Adversarial Vulnerabilities in Neural Operator Digital Twins: Gradient-Free Attacks on Nuclear Thermal-Hydraulic Surrogates

  • 用无梯度进化算法发现模型对边界条件极度敏感,仅改1%输入即引发灾难性错误。
  • 攻击使相对误差从1.5%飙升至37%-63%,且100%躲过异常检测。
  • 提出有效扰动维度,解释为何部分高敏感模型反而不易被攻破。

神经算子模型正成为核能与能源系统数字孪生的预测核心,可从稀疏传感器数据实时重构场域。然而其对抗扰动下的鲁棒性尚未明确,这对安全关键系统部署构成重大隐患。本文通过无梯度差分进化方法,在四种算子架构上验证:极稀疏(少于1%输入)且物理解释合理的扰动,能利用模型对边界条件的敏感性,触发灾难性预测失败。相对 $L_2$ 误差从约1.5%(验证精度)跃升至37%-63%,而标准验证指标完全无法察觉。值得注意的是,所有成功单点攻击均通过z-score异常检测。我们引入有效扰动维度 $d_{ ext{eff}}$,结合敏感度大小,建立双因素脆弱性模型:极端敏感集中(如POD-DeepONet,$d_{ ext{eff}} \approx 1$)因低秩输出投影限制最大误差,反而不最易被攻破;而中等敏感集中但具足够放大能力(如S-DeepONet,$d_{ ext{eff}} \approx 4$)产生最高攻击成功率。无梯度搜索在具有梯度路径问题的架构上优于梯度基方法(PGD),同等幅度随机扰动成功率接近零,证实漏洞为结构性。研究揭示了算子学习模型此前被忽视的攻击面,表明此类模型部署前必须具备超越常规验证的鲁棒性保障。

原文摘要 · Abstract (English)

Operator learning models are rapidly emerging as the predictive core of digital twins for nuclear and energy systems, promising real-time field reconstruction from sparse sensor measurements. Yet their robustness to adversarial perturbations remains uncharacterized, a critical gap for deployment in safety-critical systems. Here we show that neural operators are acutely vulnerable to extremely sparse (fewer than 1% of inputs), physically plausible perturbations that exploit their sensitivity to boundary conditions. Using gradient-free differential evolution across four operator architectures, we demonstrate that minimal modifications trigger catastrophic prediction failures, increasing relative $L_2$ error from $\sim$1.5% (validated accuracy) to 37-63% while remaining completely undetectable by standard validation metrics. Notably, 100% of successful single-point attacks pass z-score anomaly detection. We introduce the effective perturbation dimension $d_{\text{eff}}$, a Jacobian-based diagnostic that, together with sensitivity magnitude, yields a two-factor vulnerability model explaining why architectures with extreme sensitivity concentration (POD-DeepONet, $d_{\text{eff}} \approx 1$) are not necessarily the most exploitable, since low-rank output projections cap maximum error, while moderate concentration with sufficient amplification (S-DeepONet, $d_{\text{eff}} \approx 4$) produces the highest attack success. Gradient-free search outperforms gradient-based alternatives (PGD) on architectures with gradient pathologies, while random perturbations of equal magnitude achieve near-zero success rates, confirming that the discovered vulnerabilities are structural. Our findings expose a previously overlooked attack surface in operator learning models and establish that these models require robustness guarantees beyond standard validation before deployment.

神经算子数字孪生对抗攻击核能系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。