arXiv:2506.12322cs.LGcs.AI2025-06综述被引 19

针对生物制药小数据难题,系统梳理了实用机器学习方法。

Machine Learning Methods for Small Data and Upstream Bioprocessing Applications: A Comprehensive Review

  • 按应对小数据的思路分类机器学习方法
  • 分析各类方法在生物工艺中的实际效果
  • 为资源受限场景提供可落地的模型选择指南

机器学习应用依赖数据,但在生物制药等资源密集型领域,获取大规模数据成本高、耗时长。上游生物工艺涉及活细胞培养与优化以生产治疗性蛋白和生物制剂,其复杂性与高资源需求常导致数据量有限。本文全面综述了专为小数据设计的机器学习方法,构建分类体系以指导实际应用。对每类方法深入分析其核心原理,并评估其在上游生物工艺及其他相关领域中应对小数据挑战的有效性。通过多角度剖析,本综述提供了可操作的见解,识别研究空白,并为数据受限环境下的机器学习应用提供指引。

原文摘要 · Abstract (English)

Data is crucial for machine learning (ML) applications, yet acquiring large datasets can be costly and time-consuming, especially in complex, resource-intensive fields like biopharmaceuticals. A key process in this industry is upstream bioprocessing, where living cells are cultivated and optimised to produce therapeutic proteins and biologics. The intricate nature of these processes, combined with high resource demands, often limits data collection, resulting in smaller datasets. This comprehensive review explores ML methods designed to address the challenges posed by small data and classifies them into a taxonomy to guide practical applications. Furthermore, each method in the taxonomy was thoroughly analysed, with a detailed discussion of its core concepts and an evaluation of its effectiveness in tackling small data challenges, as demonstrated by application results in the upstream bioprocessing and other related domains. By analysing how these methods tackle small data challenges from different perspectives, this review provides actionable insights, identifies current research gaps, and offers guidance for leveraging ML in data-constrained environments.

机器学习小数据生物工艺综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。