用信息论指导数据采样,提升复杂系统建模效率。
Information theory and discriminative sampling for model discovery
- 基于费舍尔信息矩阵优化数据采样策略。
- 在单轨迹、可调参数等场景下显著降低数据需求。
- 适合需要高效建模的物理系统发现任务。
费舍尔信息与香农熵是从互补角度理解动力系统的核心工具。它们能通过量化变量中包含的信息来表征未知参数,或评估不同初始轨迹或时间片段对系统动力学学习或推断的贡献。本文将费舍尔信息矩阵(FIM)引入数据驱动的非线性动力系统稀疏识别(SINDy)框架。通过对混沌与非混沌系统的单条及多初始条件轨迹进行可视化分析,展示了信息分析如何通过优先选择更富信息量的数据来提升采样效率并增强模型性能。谱分析进一步揭示了统计袋装(bagging)的优势。研究还说明,在仅有一条轨迹、存在可调控制参数、或可自由初始化多条轨迹三种情形下,利用费舍尔信息与熵度量均可促进数据效率。随着数据驱动模型发现日益重要,基于可量化信息度量的合理采样策略,为提升学习效率、减少数据依赖提供了有力途径。
原文摘要 · Abstract (English)
Fisher information and Shannon entropy are fundamental tools for understanding and analyzing dynamical systems from complementary perspectives. They can characterize unknown parameters by quantifying the information contained in variables, or measure how different initial trajectories or temporal segments of a trajectory contribute to learning or inferring system dynamics. In this work, we leverage the Fisher Information Matrix (FIM) within the data-driven framework of {\em sparse identification of nonlinear dynamics} (SINDy). We visualize information patterns in chaotic and non-chaotic systems for both single trajectories and multiple initial conditions, demonstrating how information-based analysis can improve sampling efficiency and enhance model performance by prioritizing more informative data. The benefits of statistical bagging are further elucidated through spectral analysis of the FIM. We also illustrate how Fisher information and entropy metrics can promote data efficiency in three scenarios: when only a single trajectory is available, when a tunable control parameter exists, and when multiple trajectories can be freely initialized. As data-driven model discovery continues to gain prominence, principled sampling strategies guided by quantifiable information metrics offer a powerful approach for improving learning efficiency and reducing data requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。