arXiv:2606.19245cs.AIcs.LG2026-06被引 1

首个针对小分子药物早期药理学的AI评估基准,测试真实数据推理能力。

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

论文配图:TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
图 1 · 摘自论文原文
  • 构建真实药理数据任务,要求AI在编码环境中分析实验文件并输出结构化结论。
  • 16种模型配置中最强表现仅达59.3%正确率,多数系统无法稳定完成决策。
  • 适合关注AI在药物发现中实际应用潜力的研究者与开发者参考。

人工智能代理有望通过压缩解读与决策循环加速药物发现,但实际部署需可信评估。本文提出治疗学基准-早期药理学(TxBench-PP),一个可验证的小分子早期药理学基准,是更广泛治疗学基准计划的第一阶段。该基准测试代理是否能从真实实验数据中恢复准确结论,而非依赖文献记忆。包含100项评估,覆盖药物作用机制(MoA)、药效学(PD)推理、化合物-靶点结合、因果靶点验证、可成药性与安全性、转化有效性等任务,按项目阶段、检测类型和任务结构划分。代理接收真实工作流快照,在代码环境中检查文件,返回结构化答案并由确定性标准评分。在16种模型-配置组合(11个模型,4800条轨迹)中,无一系统能可靠完成药理学决策。最强配置Claude Opus 4.8 / Pi在300次终点尝试中成功178次(59.3%;95% CI: 51.1–67.6),其次为GPT-5.5 / Pi(166/300,55.3%;47.0–63.6)。

原文摘要 · Abstract (English)

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

AI药物发现药理评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。