arXiv:2410.19631cs.LG2024-10ICLR被引 2

通过智能筛选难例降低生物实验成本

Efficient Biological Data Acquisition through Inference Set Design

  • 基于置信度主动学习,优先标注最难样本
  • 实验成本降低60%以上,性能仍保持高位
  • 适合药物研发中高通量筛选场景

在药物发现中,自动化高通量实验室需测试大量化合物,但实验成本高昂。本文将此问题建模为序列子集选择任务:在保证系统整体达到目标精度的前提下,选择最少数量的化合物进行实验,其余通过预测获得结果。关键观察是,若输入空间中预测难度存在异质性,则优先获取难样本的标签后,剩余待预测样本将主要为较易样本,从而提升整体性能。该机制称为推理集设计,我们提出基于置信度的主动学习方法来剔除挑战性样本,并引入显式停止条件,在系统足够自信达到目标性能时终止采集循环。在图像数据集、分子数据集及真实世界大规模生物检测实验上的实证研究显示,该方法显著降低实验成本,同时维持高系统性能。

原文摘要 · Abstract (English)

In drug discovery, highly automated high-throughput laboratories are used to screen a large number of compounds in search of effective drugs. These experiments are expensive, so one might hope to reduce their cost by only experimenting on a subset of the compounds, and predicting the outcomes of the remaining experiments. In this work, we model this scenario as a sequential subset selection problem: we aim to select the smallest set of candidates in order to achieve some desired level of accuracy for the system as a whole. Our key observation is that, if there is heterogeneity in the difficulty of the prediction problem across the input space, selectively obtaining the labels for the hardest examples in the acquisition pool will leave only the relatively easy examples to remain in the inference set, leading to better overall system performance. We call this mechanism inference set design, and propose the use of a confidence-based active learning solution to prune out these challenging examples. Our algorithm includes an explicit stopping criterion that interrupts the acquisition loop when it is sufficiently confident that the system has reached the target performance. Our empirical studies on image and molecular datasets, as well as a real-world large-scale biological assay, show that active learning for inference set design leads to significant reduction in experimental cost while retaining high system performance.

主动学习药物发现实验优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。