用微生物生长方程知识提升数据稀缺下的生物制造建模效果
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling

- 用模拟的生物动力学方程数据预训练神经网络解码器
- 在11个数据集、7种微生物上均显著优于无先验基线
- 模拟预训练可替代真实生物结构设计,适合资源有限的研究
深度学习虽加速了药物发现,但在生物制造领域影响有限,主因是数据稀缺。生物反应器实验成本高、耗时数天至数周,且极少公开,导致每项研究仅有少量实验数据。然而,该领域蕴含丰富先验知识:生物动力学常微分方程(ODE)模型已描述微生物生长数十年,但如何将此类知识注入神经网络尚未系统研究。本文首次系统探讨将ODE知识注入神经网络的方法,比较了数据级先验(在模拟ODE曲线数据上预训练通用解码器)与架构级先验(在解码器中嵌入真实生物结构)。两者在11个数据集和7种微生物上均一致优于无先验基线。核心发现为二者可互换:在模拟数据上预训练的通用解码器表现等同于在真实数据上训练的全生物结构解码器。因此,模拟预训练提供了一种简单、高效的数据利用方案,适用于生物过程数据稀缺场景。
原文摘要 · Abstract (English)
While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data scarcity. Bioreactor experiments are high-cost, take days to weeks, and are rarely shared in public form, leaving each research work with only a handful of experiments. The domain itself, however, is rich in prior knowledge. Biokinetic ordinary differential equation (ODE) models have described microbial growth for decades, yet how to inject this knowledge into a neural network has not been studied systematically. We present the first systematic study of how to inject this ODE knowledge into a neural network, comparing a data-level prior that pre-trains a generic decoder on simulated ODE curves against an architecture-level prior that embeds the ODE inside the decoder. Both consistently outperform no-prior baselines across 11 datasets and 7 microbial species. Our central finding is that the two are substitutable. A generic decoder pre-trained on simulation matches a fully bio-structured decoder trained on real data. Simulation pre-training therefore offers a simple, data-efficient recipe for deep learning under bioprocess data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。