提出基于得分筛选的字典选择方法,提升系统辨识中稀疏回归的准确性和可解释性。
From STLS to Projection-based Dictionary Selection in Sparse Regression for System Identification
- 通过得分引导筛选字典项,优化稀疏回归中的特征选择
- 在常微分方程和偏微分方程上显著提升建模精度与可读性
- 适用于追求高鲁棒性的数据驱动系统建模研究者
本文重新审视基于字典的稀疏回归,特别是序列阈值最小二乘法(STLS),提出一种基于得分的库筛选策略,为数据驱动建模提供实用指导,尤其关注SINDy类算法。STLS是一种求解ℓ₀稀疏最小二乘问题的算法,通过分解高效求解最小二乘部分,并用近端方法处理稀疏项。其系数向量的成分依赖于投影重构误差(即得分)以及字典项间的互相干性。本文第一个贡献是针对得分与字典选择策略的理论分析,适用于原始及弱SINDy情形。第二,对常微分方程与偏微分方程的数值实验表明,基于得分的筛选能有效提升动态系统识别的准确性和可解释性。结果表明,在某些情况下,将得分引导方法融入字典精炼过程,有助于提升SINDy用户在发现控制方程时的鲁棒性。
原文摘要 · Abstract (English)
In this work, we revisit dictionary-based sparse regression, in particular, Sequential Threshold Least Squares (STLS), and propose a score-guided library selection to provide practical guidance for data-driven modeling, with emphasis on SINDy-type algorithms. STLS is an algorithm to solve the $\ell_0$ sparse least-squares problem, which relies on splitting to efficiently solve the least-squares portion while handling the sparse term via proximal methods. It produces coefficient vectors whose components depend on both the projected reconstruction errors, here referred to as the scores, and the mutual coherence of dictionary terms. The first contribution of this work is a theoretical analysis of the score and dictionary-selection strategy. This could be understood in both the original and weak SINDy regime. Second, numerical experiments on ordinary and partial differential equations highlight the effectiveness of score-based screening, improving both accuracy and interpretability in dynamical system identification. These results suggest that integrating score-guided methods to refine the dictionary more accurately may help SINDy users in some cases to enhance their robustness for data-driven discovery of governing equations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。