arXiv:2602.13021cs.LGcs.AI2026-02被引 1

用科学先验约束方程发现,避免伪公式陷阱。

Prior-Guided Symbolic Regression: Towards Scientific Consistency in Equation Discovery

  • 引入先验约束检查器和渐进式约束评估机制
  • 在多领域实验中显著优于现有方法,抗噪且数据稀缺下仍稳定
  • 理论证明可降低假设空间复杂度,防止伪方程产生

符号回归(SR)旨在从观测数据中发现可解释的方程,揭示自然现象背后的原理。然而,现有方法常陷入伪方程陷阱:虽拟合数据良好,却违背基本科学规律。根本原因在于其依赖经验风险最小化,缺乏显式约束保障科学一致性。为此,我们提出PG-SR框架,基于三阶段流程——预热、演化与精炼,全程引入先验约束检查器,将领域先验编码为可执行的约束程序,并在演化阶段采用先验退火约束评估(PACE)机制,逐步引导发现过程向科学一致区域收敛。理论上,我们证明了PG-SR能降低假设空间的Rademacher复杂度,获得更紧的泛化界,从而提供对伪方程的保证。实验表明,PG-SR在多个领域超越当前最优基线,在先验质量变化、噪声数据及数据稀缺条件下均保持鲁棒性。

原文摘要 · Abstract (English)

Symbolic Regression (SR) aims to discover interpretable equations from observational data, with the potential to reveal underlying principles behind natural phenomena. However, existing approaches often fall into the Pseudo-Equation Trap: producing equations that fit observations well but remain inconsistent with fundamental scientific principles. A key reason is that these approaches are dominated by empirical risk minimization, lacking explicit constraints to ensure scientific consistency. To bridge this gap, we propose PG-SR, a prior-guided SR framework built upon a three-stage pipeline consisting of warm-up, evolution, and refinement. Throughout the pipeline, PG-SR introduces a prior constraint checker that explicitly encodes domain priors as executable constraint programs, and employs a Prior Annealing Constrained Evaluation (PACE) mechanism during the evolution stage to progressively steer discovery toward scientifically consistent regions. Theoretically, we prove that PG-SR reduces the Rademacher complexity of the hypothesis space, yielding tighter generalization bounds and establishing a guarantee against pseudo-equations. Experimentally, PG-SR outperforms state-of-the-art baselines across diverse domains, maintaining robustness to varying prior quality, noisy data, and data scarcity.

符号回归科学发现先验约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。