提出新方法让非线性动力学识别摆脱数据缩放干扰,提升模型可靠性。
Towards a data-scale independent regulariser for robust sparse identification of non-linear dynamics
- 用统计显著性替代大小阈值,避免归一化导致的误判
- 在归一化噪声数据上,准确率远超传统STLSQ和E-SINDy
- 适合物理建模、工程系统分析等需可解释模型的场景
数据归一化在工程与科学应用中常见,但会严重扭曲基于幅度的稀疏回归方法对控制方程的发现。这一问题在稀疏非线性动力学识别(SINDy)框架中尤为突出,其核心稀疏性假设因数据缩放与测量噪声的交互而被破坏,导致发现的模型稠密、不可解释且物理错误。为解决此关键缺陷,本文提出序列系数变异阈值法(STCV),一种计算高效、天然抗数据缩放的稀疏回归算法。STCV以无量纲统计指标‘系数存在度’(CP)替代传统幅度阈值,评估模型库中候选项的统计有效性与一致性。该转变使发现过程对任意数据缩放保持不变。通过在典型动力系统及实际工程问题(包括物理弹簧-质量-阻尼实验)上的全面基准测试,结果表明,STCV在归一化、含噪数据上始终显著优于标准序列阈值最小二乘法(STLSQ)和集成SINDy(E-SINDy)。STCV方法能在其他方法失效时成功识别正确、稀疏的物理定律。通过消除归一化的扭曲效应,STCV使稀疏系统识别更可靠、自动化,增强模型可解释性与可信度。
原文摘要 · Abstract (English)
Data normalisation, a common and often necessary preprocessing step in engineering and scientific applications, can severely distort the discovery of governing equations by magnitudebased sparse regression methods. This issue is particularly acute for the Sparse Identification of Nonlinear Dynamics (SINDy) framework, where the core assumption of sparsity is undermined by the interaction between data scaling and measurement noise. The resulting discovered models can be dense, uninterpretable, and physically incorrect. To address this critical vulnerability, we introduce the Sequential Thresholding of Coefficient of Variation (STCV), a novel, computationally efficient sparse regression algorithm that is inherently robust to data scaling. STCV replaces conventional magnitude-based thresholding with a dimensionless statistical metric, the Coefficient Presence (CP), which assesses the statistical validity and consistency of candidate terms in the model library. This shift from magnitude to statistical significance makes the discovery process invariant to arbitrary data scaling. Through comprehensive benchmarking on canonical dynamical systems and practical engineering problems, including a physical mass-spring-damper experiment, we demonstrate that STCV consistently and significantly outperforms standard Sequential Thresholding Least Squares (STLSQ) and Ensemble-SINDy (E-SINDy) on normalised, noisy datasets. The results show that STCV-based methods can successfully identify the correct, sparse physical laws even when other methods fail. By mitigating the distorting effects of normalisation, STCV makes sparse system identification a more reliable and automated tool for real-world applications, thereby enhancing model interpretability and trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。