现有生物序列设计评估方法因依赖不一致的预测模型,结果不可靠。
Overconfident Oracles: Limitations of In Silico Sequence Design Benchmarking
- 用不同架构或训练种子的预言模型评估序列设计,导致性能对比失真。
- 同一方法在不同预言模型下表现差异大,反映模型泛化能力差。
- 建议加入生物物理指标,提升生成序列的可行性评估可靠性。
机器学习可自动化设计生物序列,以降低实验成本并加速医学研究。由于难以访问湿实验平台,当前方法普遍使用预言模型(oracle)来评估新生成序列。然而,不同方法采用的预言模型各异,使得比较变得不可靠。本文分析了12种使用常见机器学习预言模型的序列设计方法,发现其跨一致性与可复现性存在显著问题。不同架构甚至仅训练种子差异即导致性能排名冲突,表明模型在分布外数据上泛化能力差是核心问题。为此,我们提出补充一系列生物物理度量,用于评估生成序列的可行性,并限制预言模型需评分的分布范围,从而增强设计流程的鲁棒性。本工作旨在揭示当前评估流程的潜在陷阱,推动更稳健基准的建立,最终促进体外序列设计方法的改进。
原文摘要 · Abstract (English)
Machine learning methods can automate the in silico design of biological sequences, aiming to reduce costs and accelerate medical research. Given the limited access to wet labs, in silico design methods commonly use an oracle model to evaluate de novo generated sequences. However, the use of different oracle models across methods makes it challenging to compare them reliably, motivating the question: are in silico sequence design benchmarks reliable? In this work, we examine 12 sequence design methods that utilise ML oracles common in the literature and find that there are significant challenges with their cross-consistency and reproducibility. Indeed, oracles differing by architecture, or even just training seed, are shown to yield conflicting relative performance with our analysis suggesting poor out-of-distribution generalisation as a key issue. To address these challenges, we propose supplementing the evaluation with a suite of biophysical measures to assess the viability of generated sequences and limit out-of-distribution sequences the oracle is required to score, thereby improving the robustness of the design procedure. Our work aims to highlight potential pitfalls in the current evaluation process and contribute to the development of robust benchmarks, ultimately driving the improvement of in silico design methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。