arXiv:2608.23075astro-ph.IMastro-ph.HE2026-08

神经代理加速物理模拟失效,因运行时占比低、分布外检测无效、误差累积难控。

When a neural surrogate cannot accelerate a solver: runtime share, closed-loop drift, and the economics of uncertainty gating in a stiff coupled simulation

论文配图:When a neural surrogate cannot accelerate a solver: runtime share, closed-loop drift, and the economics of uncertainty gating in a stiff coupled simulation
图 1 · 摘自论文原文
  • 代理模型虽单次计算快5.8倍,但占总时间仅16.9%,极限加速不足1.2倍
  • 训练精度无法预测部署效果,真实场景中误差与存活率相关性为负(rho=-0.04)
  • 分布外检测会延迟96.8%以上计算单元,导致整体变慢0.94~0.96倍

针对高成本的隐式牛顿求解器,研究发现基于学习的代理模型难以实现加速。关键障碍有三:第一,目标模块占总运行时仅16.9%,按阿姆达尔定律极限加速不超过1.2倍;即使代理模型单次计算快5.8倍,也仅能达成性能持平。第二,离线精度无法衡量部署有效性——14个模型中,误差与存活率的总体相关性(rho=+0.73)被家族因素干扰,控制后降为负相关(rho=-0.04)。第三,分布外门控机制在闭环系统中无效:因状态偏离数据流形73倍,门控需延迟96.8%至99.7%的网格单元,计入自身开销后,整体运行效率下降至0.94~0.96倍。此外,即使不崩溃的运行也积累线性-19.9%密度偏差,为定向累积而非随机发散。

原文摘要 · Abstract (English)

Learned surrogates for expensive inner solver blocks are a widely pursued route to faster multiphysics simulation. We report a controlled, end-to-end negative result and identify three structural barriers, none of them a deficiency of the network we trained. The testbed is the implicit Newton solve coupling energy-dependent neutrino radiation to matter in a general-relativistic radiation-hydrodynamics code, its most expensive physics routine per call. First, per-call cost and share of runtime are different quantities, and only the second bounds acceleration. An exclusive self-time profile puts the target block at 16.9% of critical-rank wall clock, capping any surrogate at ~1.2x by Amdahl's law. A surrogate 5.8x cheaper per call merely ties the solver, and the configuration stable enough to run without fallback reaches only parity. Second, offline accuracy cannot rank surrogates for deployment: across fourteen networks the pooled Spearman error-versus-survival correlation (rho = +0.73) is a between-family confound that vanishes under control (rho = -0.04). Third, a correct out-of-distribution gate cannot accelerate a loop that leaves its training distribution. We give the break-even deferral fraction in closed form: because the visited states sit 73x off the data manifold, the gate defers 96.8 to 99.7% of cells, almost invariant to surrogate quality. Including its own cost, the gated loop is a 0.94 to 0.96x slowdown. We further separate stability from fidelity: a never-crashing gated run accumulates a linear -19.9% density bias over 6000 steps. The error is a directed, ballistically accumulating bias, not the variance-driven divergence the autoregressive literature targets.

神经代理物理模拟误差累积分布外

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。