arXiv:2603.11330q-bio.QMcs.LG2026-03被引 2

研究生物系统方程学习中的数值病态问题,揭示其对模型准确性的影响。

Ill-Conditioning in Dictionary-Based Dynamic-Equation Learning: A Systems Biology Case Study

  • 通过系统分析发现,仅两三个项组合就可能引发严重多重共线性。
  • 数据分布偏离正交基权重函数时,正交多项式反而比单项式更差。
  • 数据与正交基匹配时,模型恢复精度显著提升,适合生物动力学建模研究者。

从时间序列数据中进行数据驱动的控制方程发现为理解复杂生物系统提供了有力框架。基于库的方法通过候选函数的稀疏回归表现出巨大潜力,但在候选函数高度相关时面临关键挑战:数值病态。采样不足或特定候选库选择可导致强多重共线性和数值不稳定性。此时,测量噪声可能导致截然不同的恢复模型,掩盖真实动态并阻碍准确系统识别。尽管稀疏正则化可促进简约解并部分缓解条件问题,但强相关性仍可能存在,正则化可能引入偏差,且回归问题对数据微小扰动仍高度敏感。本文针对系统生物学基准模型,系统分析了病态性对稀疏生物动态识别的影响。结果显示,仅两个或三个项的组合即可产生强多重共线性和极高的条件数。进一步表明,正交多项式基并不总能解决病态问题,在数据分布偏离其对应权重函数时,性能甚至劣于单项式库。最后,当数据采样分布与正交基的适当权重函数对齐时,数值条件改善,正交多项式基在两个基准模型上均提升了模型恢复精度。

原文摘要 · Abstract (English)

Data-driven discovery of governing equations from time-series data provides a powerful framework for understanding complex biological systems. Library-based approaches that use sparse regression over candidate functions have shown considerable promise, but they face a critical challenge when candidate functions become strongly correlated: numerical ill-conditioning. Poor or restricted sampling, together with particular choices of candidate libraries, can produce strong multicollinearity and numerical instability. In such cases, measurement noise may lead to widely different recovered models, obscuring the true underlying dynamics and hindering accurate system identification. Although sparse regularization promotes parsimonious solutions and can partially mitigate conditioning issues, strong correlations may persist, regularization may bias the recovered models, and the regression problem may remain highly sensitive to small perturbations in the data. We present a systematic analysis of how ill-conditioning affects sparse identification of biological dynamics using benchmark models from systems biology. We show that combinations involving as few as two or three terms can already exhibit strong multicollinearity and extremely large condition numbers. We further show that orthogonal polynomial bases do not consistently resolve ill-conditioning and can perform worse than monomial libraries when the data distribution deviates from the weight function associated with the orthogonal basis. Finally, we demonstrate that when data are sampled from distributions aligned with the appropriate weight functions corresponding to the orthogonal basis, numerical conditioning improves, and orthogonal polynomial bases can yield improved model recovery accuracy across two baseline models.

系统生物学方程发现数值病态稀疏回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。