arXiv:2606.02624q-bio.QMcs.AI2026-06中稿 · ICML

构建百万级蛋白质进化数据集,评估AI预测未来实验结果的能力

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

论文配图:TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering
图 1 · 摘自论文原文
  • 基于31轮定向进化实验构建可复现的未来轮次预测基准
  • 模型在跨轮次排序任务中表现远低于插值预测,凸显泛化挑战
  • 适合研究智能蛋白质工程系统的算法与数据设计

AI驱动科学发现正步入代理时代,蛋白质工程系统需优先预测未来湿实验结果而非仅拟合静态数据。我们提出TadA-Bench,一个来自31轮TadA定向进化实验的百万变体湿实验重播基准,用于未来轮次发现的代理式蛋白质工程研究。该基准保留实验进程时间线,定义固定数据重播任务:基于前期实验轮次,对仅在后期出现的变异体进行排序。数据提供对齐的DNA、RNA和蛋白质视图,并采用Seq2Graph——一种基于图的标签统一流程——将嘈杂的富集测量值整合为跨轮次一致的活性标签。随机分割对照显示模型具备强内插能力,但未来轮次排序与有限预算候选选择性能显著更弱。受控分析表明,进化覆盖范围比局部数据密度更具信息量,使TadA-Bench成为可复现的湿实验重播基础平台,助力未来轮次发现的代理式蛋白质工程;数据与代码已发布于Hugging Face和GitHub。

原文摘要 · Abstract (English)

AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather than merely fit static measurements. We introduce TadA-Bench, a million-variant wet-lab replay benchmark from 31 TadA directed-evolution rounds for future-round discovery toward agentic protein engineering. TadA-Bench preserves the campaign chronology and defines a fixed-data replay task: given earlier experimental rounds, models rank variants that appear only in later rounds. It provides aligned DNA, RNA, and protein views, and uses Seq2Graph, a graph-based label-unification pipeline, to reconcile noisy enrichment measurements into consistent cross-round activity labels. Random-split controls show strong interpolation, but future-round ranking and finite-budget candidate selection are much weaker. Controlled analyses suggest that evolutionary coverage is more informative than local data density, positioning TadA-Bench as a reproducible wet-lab replay substrate for future-round discovery toward agentic protein engineering; the data and code are released on Hugging Face and GitHub.

蛋白质工程代理智能基准测试定向进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。