arXiv:2607.18490cs.LGnlin.CD2026-07被引 2

吸引子几何决定方程发现的可识别极限,混沌未必更好。

Attractor Geometry Determines the Identifiability Limits of System Discovery

论文配图:Attractor Geometry Determines the Identifiability Limits of System Discovery
图 1 · 摘自论文原文
  • 用洛伦兹-84系统验证:吸引子覆盖函数空间程度由最小特征值λ_min(M)决定
  • λ_min(M)越小恢复越难,混沌虽提升其值但放大噪声,影响不同算法表现
  • 提出软F1指标,揭示传统评分忽略的结构性能差异,适合机制发现研究者

从数据中进行符号方程发现不仅受限于算法设计和数据量,更受吸引子几何制约——即长期动力学允许恢复的内容。在洛伦兹-84系统中,单一扰动参数驱动固定点、周期轨和混沌态,而方程与候选库保持不变。我们发现,一个数值 $λ_{ ext{min}}(M)$(不变测度矩矩阵的最小特征值)决定了稀疏回归(SINDy)与进化符号回归(PySR)的可识别上限。该值源自Birkhoff遍历定理,仅需短参考轨迹即可计算,反映吸引子对函数空间的覆盖程度:当其趋近零时,任何算法均无法恢复;随其增大,两类方法性能同步提升。混沌通过扩展吸引子提高 $λ_{ ext{min}}(M)$,但也扩大其范围并增强噪声;由于噪声在SINDy中线性进入回归瓶颈,而在PySR中超线性进入判别通道,相同转变可能使两方法表现相反,故更深混沌并非总有利。此框架下的无参机制评分可无需重训练迁移至洛伦兹-96系统,验证了机制而非拟合;由方程导出的准则可预测增加混沌不会改善条件的情况。此外,我们引入软F1,一种系数加权的结构化度量,揭示了二元成功与预测分数无法察觉的性能差异。因此,发现的第一问题不是选哪个算法,而是吸引子允许什么。

原文摘要 · Abstract (English)

Symbolic discovery of governing equations from data is limited not only by algorithm design and data volume, but by the geometry of the attractor: what the long-run dynamics allow to be recovered. Using a within-system design on Lorenz-84, where one forcing parameter drives fixed-point, limit-cycle, and chaotic regimes while the governing equations and library stay fixed, we show that a single number, $λ_{\min}(M)$, the smallest eigenvalue of the invariant-measure moment matrix, sets the identifiability ceiling for both sparse regression (SINDy) and evolutionary symbolic regression (PySR). Derived from the Birkhoff ergodic theorem and obtained from a short reference trajectory before any run, $λ_{\min}(M)$ measures how fully the attractor covers function space: where it vanishes, recovery is impossible for any algorithm, sparse or combinatorial alike; as it grows, both algorithms improve. Chaos raises $λ_{\min}(M)$ by spreading the attractor, but also enlarges it and amplifies noise; because noise enters SINDy's regression bottleneck linearly and PySR's discrimination channel superlinearly, the same transition can push the two methods in opposite directions, so deeper chaos is not uniformly better. Parameter-free mechanistic scores from this framework transfer without refitting to a held-out Lorenz-96 system, confirming mechanism rather than curve-fitting; a criterion read from the equations predicts when added chaos will not improve conditioning. We also introduce Soft F1, a coefficient-weighted structural metric that resolves performance differences invisible to binary-success and predictive scores. The first question of discovery is then not which algorithm, but what the attractor permits.

系统发现吸引子几何混沌可识别性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。