arXiv:2507.04121stat.MLcond-mat.stat-mech2025-07被引 1

提出新准则PASTIS,从噪声数据中自动选出最简正确随机模型。

Model selection for stochastic dynamics: a parsimonious and principled approach

  • 基于极值理论设计惩罚项,融合候选模型规模与显著性阈值。
  • 在洛伦兹、格雷-斯科特等系统上,准确率和预测能力均优于经典方法。
  • 支持高采样间隔和强噪声数据,适合真实实验场景建模。

本论文聚焦于从含噪离散时间序列中发现随机微分方程(SDE)与随机偏微分方程(SPDE)。核心挑战是从海量候选模型中选取最简正确的模型,而传统信息准则(AIC、BIC)常受限。本文提出PASTIS(Parsimonious Stochastic Inference),一种基于极值理论的新信息准则。其惩罚项 $n_/mathcal{B} \ln(n_0/p)$ 显式包含初始候选参数库大小 $n_0$、所选模型参数数 $n_/mathcal{B}$ 与显著性阈值 $p$,其中 $p$ 表示在多模型比较中误选冗余参数的概率。在洛伦兹、奥恩斯坦-乌伦贝克、洛特卡-沃尔泰拉(SDE)及格雷-斯科特(SPDE)系统上的基准测试表明,PASTIS 在精确模型识别与预测能力上均优于AIC、BIC、交叉验证(CV)和SINDy。针对实际数据中大采样间隔(Δt)或高测量噪声(σ)带来的建模困难,本文进一步构建了鲁棒变体PASTIS-Δt与PASTIS-σ,显著拓展了方法在不完美实验数据中的适用性。PASTIS提供了一套统计严谨、验证充分且实用的随机动力学建模框架。

原文摘要 · Abstract (English)

This thesis focuses on the discovery of stochastic differential equations (SDEs) and stochastic partial differential equations (SPDEs) from noisy and discrete time series. A major challenge is selecting the simplest possible correct model from vast libraries of candidate models, where standard information criteria (AIC, BIC) are often limited. We introduce PASTIS (Parsimonious Stochastic Inference), a new information criterion derived from extreme value theory. Its penalty term, $n_\mathcal{B} \ln(n_0/p)$, explicitly incorporates the size of the initial library of candidate parameters ($n_0$), the number of parameters in the considered model ($n_\mathcal{B}$), and a significance threshold ($p$). This significance threshold represents the probability of selecting a model containing more parameters than necessary when comparing many models. Benchmarks on various systems (Lorenz, Ornstein-Uhlenbeck, Lotka-Volterra for SDEs; Gray-Scott for SPDEs) demonstrate that PASTIS outperforms AIC, BIC, cross-validation (CV), and SINDy (a competing method) in terms of exact model identification and predictive capability. Furthermore, real-world data can be subject to large sampling intervals ($Δt$) or measurement noise ($σ$), which can impair model learning and selection capabilities. To address this, we have developed robust variants of PASTIS, PASTIS-$Δt$ and PASTIS-$σ$, thus extending the applicability of the approach to imperfect experimental data. PASTIS thus provides a statistically grounded, validated, and practical methodological framework for discovering simple models for processes with stochastic dynamics.

模型选择随机微分方程数据驱动建模极值理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。