首个可验证的表观基因组分析基准,测试AI能否做出科学决策。
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

- 构建106项表观基因组任务,评估AI在真实工作流中做判断的能力。
- 16个模型对318次任务仅45%通过,最高准确率不足一半。
- 模型常找对文件但缺乏专业判断,适合研究AI在生命科学中的局限。
我们提出EpiBench,一个面向短时程表观基因组分析的可验证基准。该基准评估智能体能否从真实的流程状态中做出明确的分析决策,并返回可确定性评分的答案。基准涵盖CUT&Tag/CUT&RUN、ATAC-seq、ChIP-seq和DNA甲基化四种实验流程,共106项评估。在来自16个模型-工具对的5,088条有效轨迹中,无一系统在多数任务中通过:GPT-5.5 / Pi以45.0%(143/318次;95%置信区间,36.3–53.7)表现最佳,其次为GPT-5.5 / OpenAI Codex(39.9%,127/318次;95% CI,31.6–48.3)。Claude Opus 4.8 Max / Pi与GPT-5.4 / Pi均通过39.0%(124/318次;95% CI分别为30.2–47.8与31.0–47.0)。性能因检测类型而异,许多失败任务仍包含正确答案的部分内容。模型常能定位正确文件并计算出有用中间结果,但在需要深入、检测特异性科学判断的任务中失败。
原文摘要 · Abstract (English)
We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。