arXiv:2608.13504cs.LG2026-08中稿 · oral presentation …

用稀疏正交回归从噪声数据中自动发现微分方程,提升稳定性与泛化能力。

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

  • 直接通过L1正则回归估计正交基展开系数,无需数值积分或内积计算。
  • 在稀疏采样和噪声条件下仍保持稳定,优于传统稀疏回归基线。
  • 适合需要自适应基函数设计的物理系统建模与高维积分估计场景。

我们提出稀疏正交回归技术(SORT),一种用于从噪声和非均匀采样数据中学习正交基展开的稀疏谱框架。SORT通过L1-regularized回归直接从观测数据估计展开系数,避免了显式的数值求积或解析内积计算。核心应用是数据驱动的常微分方程发现:向量场在选定的正交基中表示为稀疏系数展开,提供了一种与符号回归、语法引导发现及SINDy风格稀疏识别互补的方法。该方法先恢复紧凑的谱表示,后续可指导更简单的解析形式搜索。在动力系统实验中,当基函数适配问题时,SORT性能匹配或超越基于库的稀疏回归基线,并在稀疏采样、噪声导数估计和表示不匹配下表现更稳定。具体案例表明,若有限库遗漏特定非线性,模型可能失效;而SORT虽非免疫于不匹配,但将问题从对通用项的脆弱选择,转变为针对领域定制的基函数设计。实验还显示,随着模型阶数增加,主导的低阶系数持续存在,支持有序一致的模型增长。除方程发现外,同一学习展开还可用于非线性逼近与复杂高维积分的系数读出。总体上,SORT为系统识别、逼近与积分提供可复用的中间表示,同时将基函数设计明确纳入科学建模过程。

原文摘要 · Abstract (English)

We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation. The central application is data-driven discovery of ordinary differential equations: vector fields are represented in chosen orthogonal bases and learned as sparse coefficient expansions. This provides a complementary route to symbolic regression, grammar-based discovery, and SINDy-style sparse identification by first recovering a compact spectral representation, which can later guide searches for simpler analytic forms. Across the dynamical-system experiments, SORT matches or improves upon library-based sparse-regression baselines when the basis is well adapted to the problem, and shows more stable degradation under sparse sampling, noisy derivative estimates, and representation mismatch. Specific examples illustrate why this representation is useful: if a finite library misses the problem-specific nonlinearity, the resulting model can fail. SORT is not immune to mismatch, but it shifts the problem away from brittle selection among generic terms to basis design adapted to the problem domain. The experiments also show that dominant low-order coefficients persist as model order increases, supporting order-consistent model growth. Beyond equation discovery, the same learned expansion supports nonlinear approximation and estimation of complex, high-dimensional integrals by coefficient readout. Overall, SORT provides a reusable intermediate representation for system identification, approximation, and integration, while making basis design an explicit part of the scientific modeling problem.

方程发现稀疏回归谱方法系统识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。