arXiv:2604.01335cs.CEcs.LG2026-04

神经符号框架揭示了模型偏差继承问题,强调真实物理验证比拟合误差更重要。

Bias Inheritance in Neural-Symbolic Discovery of Constitutive Closures Under Function-Class Mismatch

  • 三阶段框架:约束学习、符号压缩、未见初始条件重仿真验证
  • 函数类不匹配时神经模型可压缩为简洁符号表达,但误差几乎不变
  • 符号压缩无法修复原始偏差,前向验证是判断物理准确性的关键

我们研究在已知偏微分方程结构的非线性反应-扩散系统中,从时空观测数据中驱动发现本构闭包。目标是稳健恢复扩散与反应律,避免将低残差或短时预测误当作物理规律恢复。提出三阶段神经符号框架:(1) 使用噪声鲁棒的弱形式驱动目标,在物理约束下学习数值代理;(2) 将这些代理压缩为受限可解释符号族(如多项式、有理式、饱和形式);(3) 在未见初值条件下通过显式前向重仿真验证符号闭包。大量数值实验揭示两种不同情形:在函数类匹配设置下,弱多项式基线表现如同正确设定的参考估计器,表明神经代理并不总优于经典基;而在函数类不匹配时,神经代理提供必要灵活性,可压缩为紧凑符号律且滚动预测退化极小。然而,我们识别出关键的“偏差继承”机制:符号压缩无法自动修复本构偏差。在多种观测条件下,符号闭包的真实误差紧密追踪神经代理误差,偏差继承比接近1。这表明神经符号建模的主要瓶颈在于初始数值逆问题,而非后续符号压缩。强调本构声明必须由前向验证支持,而非仅依赖残差最小化。

原文摘要 · Abstract (English)

We investigate the data-driven discovery of constitutive closures in nonlinear reaction-diffusion systems with known governing PDE structures. Our objective is to robustly recover diffusion and reaction laws from spatiotemporal observations while avoiding the common pitfall where low residuals or short-horizon predictions are conflated with physical recovery. We propose a three-stage neural-symbolic framework: (1) learning numerical surrogates under physical constraints using a noise-robust weak-form-driven objective; (2) compressing these surrogates into restricted interpretable symbolic families (e.g., polynomial, rational, and saturation forms); and (3) validating the symbolic closures through explicit forward re-simulation on unseen initial conditions. Extensive numerical experiments reveal two distinct regimes. Under matched-library settings, weak polynomial baselines behave as correctly specified reference estimators, showing that neural surrogates do not uniformly outperform classical bases. Conversely, under function-class mismatch, neural surrogates provide necessary flexibility and can be compressed into compact symbolic laws with minimal rollout degradation. However, we identify a critical "bias inheritance" mechanism where symbolic compression does not automatically repair constitutive bias. Across various observation regimes, the true error of the symbolic closure closely tracks that of the neural surrogate, yielding a bias inheritance ratio near one. These findings demonstrate that the primary bottleneck in neural-symbolic modeling lies in the initial numerical inverse problem rather than the subsequent symbolic compression. We underscore that constitutive claims must be rigorously supported by forward validation rather than residual minimization alone.

神经符号反问题物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。