arXiv:2605.26631stat.APcs.LG2026-05被引 2

从噪声数据中精准识别微分方程,避免误选干扰项。

Data-driven sparse identification of governing PDEs via knockoff filters and multi-criteria trade-offs

论文配图:Data-driven sparse identification of governing PDEs via knockoff filters and multi-criteria trade-offs
图 1 · 摘自论文原文
  • 用敲除过滤器筛选候选项,控制假阳性率。
  • 在严重噪声下准确恢复真实方程结构,误差小。
  • 适合需要高可靠性建模的科研与工程场景。

我们提出KO-PDE-IDENT,一种数据驱动框架,用于在控制错误发现率(FDR)的前提下,识别简洁的偏微分方程(PDE)。从含噪观测中发现PDE常受候选项间极端多重共线性困扰,导致常规稀疏回归方法误选虚假项。KO-PDE-IDENT首先通过模型-X敲除过滤器(有限样本下FDR控制)挖掘潜在候选项集合,再精炼并排序剩余的PDE备选方案。该框架包含三部分:第一,结合ℓ₀约束自适应最优子集选择与SHAP,构建敲除特征统计量,生成高效且计算快速的差异统计量;第二,采用递归特征消除(RFE)移除边际贡献可忽略的项,并通过敲除扰动假设检验评估统计必要性;第三,将最终模型选择建模为多准则决策问题,最优控制方程是预测精度、模型复杂度与系数不确定性等多指标平衡的最佳解。我们在五类经典PDE上测试该框架,结果表明其可在严重噪声条件下精确恢复真实方程结构,完全消除假发现,保留所有真实项,且系数估计误差低。

原文摘要 · Abstract (English)

We propose KO-PDE-IDENT, a data-driven framework for identifying parsimonious partial differential equations (PDEs) with false discovery rate (FDR) control. PDE discovery from noisy observations is often hindered by extreme multicollinearity among candidate terms, which causes typical sparse-regression methods to select spurious terms. To address this problem, KO-PDE-IDENT initially mines a support set of potential candidate terms via model-X knockoff filters with finite-sample FDR control, then refines and ranks the surviving PDE alternatives. The framework integrates three components. First, knockoff feature statistics are constructed by coupling $\ell_{0}$-constrained adaptive best-subset selection with SHapley Additive exPlanations (SHAP), yielding an effective and computationally efficient difference statistic. Second, a recursive feature elimination (RFE) procedure removes terms whose marginal contributions are dispensable and assesses statistical necessity through knockoff-perturbed hypothesis testing. Third, the final model selection is formulated as a multi-criteria decision-making (MCDM) problem, where the optimal governing equation is the alternative that best balances a wide range of criteria such as predictive accuracy, model complexity and coefficient uncertainty. We evaluate KO-PDE-IDENT on five canonical PDEs under severe noise corruption. Empirical results show that our framework can exactly recover the true PDE structure, eliminating false discoveries while retaining all true underlying terms, with low coefficient estimation error.

PDE识别敲除过滤稀疏回归噪声建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。