自动构建生物医学数字孪生,提升药物研发效率
Data-driven Discovery of Digital Twins in Biomedical Research
- 用稀疏回归与贝叶斯框架从时序数据中推断生物反应网络
- 在噪声数据和高维条件下,稀疏回归优于符号回归
- 适合从事系统生物学与个性化医疗的科研人员参考
近期技术进步带来了大量高通量生物数据,使得构建生物系统或患者的数字孪生成为可能。这类计算工具能揭示驱动扰动或药物响应的关键反应网络,助力药物发现与个性化治疗。然而,当前方法仍依赖人工整合数据,亟需自动化手段。受物理领域数据驱动建模成功的启发,研究者尝试将类似方法应用于生物学,但面临噪声数据、多条件、先验知识融合、隐变量、高维性、未观测导数、候选库设计及不确定性量化等八大挑战。评估表明,稀疏回归(尤其结合贝叶斯框架)表现更优;深度学习与大语言模型展现出创新的先验知识整合能力,但可靠性有待提升。我们主张发展融合化学反应网络机制、贝叶斯不确定性量化与深度学习生成能力的混合模块化框架。为此,提出一个涵盖所有挑战的基准评估体系以推动方法发展。
原文摘要 · Abstract (English)
Recent technological advances have expanded the availability of high-throughput biological datasets, enabling the reliable design of digital twins of biomedical systems or patients. Such computational tools represent key reaction networks driving perturbation or drug response and can guide drug discovery and personalized therapeutics. Yet, their development still relies on laborious data integration by the human modeler, so that automated approaches are critically needed. The success of data-driven system discovery in Physics, rooted in clean datasets and well-defined governing laws, has fueled interest in applying similar techniques in Biology, which presents unique challenges. Here, we reviewed methodologies for automatically inferring digital twins from biological time series, which mostly involve symbolic or sparse regression. We evaluate algorithms according to eight biological and methodological challenges, associated to noisy/incomplete data, multiple conditions, prior knowledge integration, latent variables, high dimensionality, unobserved variable derivatives, candidate library design, and uncertainty quantification. Upon these criteria, sparse regression generally outperformed symbolic regression, particularly when using Bayesian frameworks. We further highlight the emerging role of deep learning and large language models, which enable innovative prior knowledge integration, though the reliability and consistency of such approaches must be improved. While no single method addresses all challenges, we argue that progress in learning digital twins will come from hybrid and modular frameworks combining chemical reaction network-based mechanistic grounding, Bayesian uncertainty quantification, and the generative and knowledge integration capacities of deep learning. To support their development, we further propose a benchmarking framework to evaluate methods across all challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。