用主动学习在极少量数据下精准发现复杂系统方程。
How Low Can You Go? Active Learning for Sparse Model Discovery in the Ultra-Low-Data Limit

- 基于SINDy方法,通过不确定性估计主动选择最有信息量的采样点。
- 在洛伦兹系统、贝格斯方程等场景中,所需数据比随机采样少一个数量级。
- 适合数据稀缺但需高精度建模的科研与工程领域,如物理仿真与控制。
识别复杂动力系统的基本方程仍是科学与工程中的核心挑战。传统方法依赖经验数据和启发式规则,而现代数据驱动方法更具灵活性且假设更少。然而,真实场景中数据获取往往成本高昂。本文提出一种超低数据条件下的主动学习策略,用于动力学发现。不同于随机采样,该方法迭代优先选择对模型识别最有益的区域。基于稀疏非线性动力学识别(SINDy)框架,采用集成扩展E-SINDy估计认知不确定性,指导常微分方程(ODE)与偏微分方程(PDE)的采样。针对洛伦兹系统,在不同数据预算与噪声水平下进行了全面分析;针对贝格斯方程(具陡峭激波前沿)与库拉莫托-西瓦辛斯基方程(空间结构复杂),分别考察了信息区域与非信息区域的差异。所有情况下,该方法均以远少于随机采样的数据量准确识别出主导动力学。
原文摘要 · Abstract (English)
Identifying the governing equations of complex dynamical systems remains a fundamental challenge across science and engineering. While early approaches relied on empirical data and heuristics, modern data-driven methods offer greater flexibility and fewer assumptions. However, data acquisition in real-world settings is often expensive. This work addresses this challenge by introducing an active learning strategy for dynamics discovery in the ultra-low data limit. Rather than sampling randomly, our method iteratively prioritizes regions that are most informative for model identification. This approach builds on Sparse Identification of Nonlinear Dynamics (SINDy), and utilizes an ensemble extension, E-SINDy, to estimate epistemic uncertainty and guide the sampling for both ordinary and partial differential equations (ODEs/PDEs). For ODEs, an exhaustive analysis is conducted on the Lorenz system across varying data budgets and noise levels. For PDEs, two systems with contrasting dynamical characteristics are examined: the Burgers' equation, where a sharp shock front creates a distinction between informative and uninformative regions, and the Kuramoto-Sivashinsky equation, which presents a more spatially complex sampling landscape. Across all scenarios, the proposed method accurately identifies the governing dynamics with significantly fewer data samples than random sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。