融合符号与统计学习,从有限数据中挖掘细胞培养的调控机制。
A Symbolic and Statistical Learning Framework to Discover Bioprocessing Regulatory Mechanism: Cell Culture Example
- 用随机微分方程建模过程变异性,结合生物先验构建候选调控机制。
- 通过贝叶斯联合学习参数与结构,提升在数据少时的模型准确性。
- 适合生物制造、数字孪生领域研究者,尤其关注机制发现与不确定性量化。
生物工艺机理建模对推动智能数字孪生制造至关重要,但受限于复杂的细胞内调控、系统随机行为及实验数据不足。本文提出一种符号与统计学习融合框架,用于识别关键调控机制并量化模型不确定性。通过随机微分方程刻画内在过程变异性,并基于生物知识预设候选调控机制。采用贝叶斯学习方法,联合学习动力学参数与调控结构,通过混合模型形式实现。为提升计算效率,开发了结合伴随敏感性分析的马尔可夫链-调整拉普拉斯算法以进行后验探索。相比现有先进贝叶斯推断方法,该框架在样本效率和鲁棒模型选择上表现更优。实证研究证明其能在数据有限条件下恢复缺失调控机制,并提升模型保真度。
原文摘要 · Abstract (English)
Bioprocess mechanistic modeling is essential for advancing intelligent digital twin representation of biomanufacturing, yet challenges persist due to complex intracellular regulation, stochastic system behavior, and limited experimental data. This paper introduces a symbolic and statistical learning framework to identify key regulatory mechanisms and quantify model uncertainty. Bioprocess dynamics is formulated with stochastic differential equations characterizing intrinsic process variability, with a predefined set of candidate regulatory mechanisms constructed from biological knowledge. A Bayesian learning approach is developed, which is based on a joint learning of kinetic parameters and regulatory structure through a formulation of the mixture model. To enhance computational efficiency, a Metropolis-adjusted Langevin algorithm with adjoint sensitivity analysis is developed for posterior exploration. Compared to state-of-the-art Bayesian inference approaches, the proposed framework achieves improved sample efficiency and robust model selection. An empirical study demonstrates its ability to recover missing regulatory mechanisms and improve model fidelity under data-limited conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。