arXiv:2510.16612stat.MLcs.LG2025-10被引 1

用生成模型优化高通量实验,大幅加速大规模生物数据学习

Accelerated Learning on Large Scale Screens using Generative Library Models

  • 仅收集活性序列正样本,结合生成模型补全负样本
  • 在10^6~10^12规模筛选中实现信息最大化,提升数据效率
  • 适合大规模蛋白质筛选与生物模型训练场景

生物机器学习常受数据规模限制。高通量筛选可并行测试10^6至10^12个蛋白序列的活性,是缓解数据瓶颈的潜在途径。本文提出算法,优化高通量筛选以高效生成数据并训练模型。针对数据量受限于测量与测序成本的大规模情形,我们证明当活性序列稀少时,仅采集活性正样本(y>0)可最大化信息增益。通过库的生成模型校正缺失的负样本,可一致且高效地估计真实p(y|x)。我们在模拟及大规模抗体筛选实验中验证该方法。整体上,实验与推断的协同设计显著加速了学习进程。

原文摘要 · Abstract (English)

Biological machine learning is often bottlenecked by a lack of scaled data. One promising route to relieving data bottlenecks is through high throughput screens, which can experimentally test the activity of $10^6-10^{12}$ protein sequences in parallel. In this article, we introduce algorithms to optimize high throughput screens for data creation and model training. We focus on the large scale regime, where dataset sizes are limited by the cost of measurement and sequencing. We show that when active sequences are rare, we maximize information gain if we only collect positive examples of active sequences, i.e. $x$ with $y>0$. We can correct for the missing negative examples using a generative model of the library, producing a consistent and efficient estimate of the true $p(y | x)$. We demonstrate this approach in simulation and on a large scale screen of antibodies. Overall, co-design of experiments and inference lets us accelerate learning dramatically.

高通量筛选生成模型数据效率生物机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。