提出可验证的科学筛选评分框架,评估AI选候选时的预算效率与准确率。
Budget-Sensitive Discovery Scoring: A Formally Verified Framework for Evaluating AI-Guided Scientific Selection
- 设计双惩罚评分机制,同时控制假发现和过度弃选,支持预算级评估
- 在HIV数据集上测试39种方法,发现传统机器学习模型优于LLM
- 框架适用于多种场景,尤其适合高成本实验筛选的公平比较
科学发现日益依赖AI系统筛选需昂贵实验验证的候选物,但缺乏兼顾预算与性能的评价框架。本文提出预算敏感发现评分(BSDS),一个经Lean 4证明助手验证的20个定理的度量标准,联合惩罚错误发现(lambda加权FDR)和过度弃选(gamma加权覆盖率缺口)。其平均形式发现质量得分(DQS)提供单一统计量,防止任何提案通过选择特定预算来夸大表现。案例研究中,在分子网HIV数据集(41,127化合物,3.5%活性,1,000次自助采样)上评估39个提案:11种机制变体、14种零样本LLM配置、14种少样本LLM配置。结果表明:简单基于随机森林的贪婪-ML提案取得最佳DQS(-0.046),优于所有多层感知机变体和所有LLM配置;无论零样本或少样本,无一LLM在HIV或Tox21上超越该基线;该排序在五个分子网基准上跨0.18%-46.2%流行率、非药物安全领域及9×7惩罚参数网格(tau≥0.636,均值0.863)中保持一致。该框架适用于任何存在预算约束与非对称误差代价的候选筛选场景。
原文摘要 · Abstract (English)
Scientific discovery increasingly relies on AI systems to select candidates for expensive experimental validation, yet no principled, budget-aware evaluation framework exists for comparing selection strategies -- a gap intensified by large language models (LLMs), which generate plausible scientific proposals without reliable downstream evaluation. We introduce the Budget-Sensitive Discovery Score (BSDS), a formally verified metric -- 20 theorems machine-checked by the Lean 4 proof assistant -- that jointly penalizes false discoveries (lambda-weighted FDR) and excessive abstention (gamma-weighted coverage gap) at each budget level. Its budget-averaged form, the Discovery Quality Score (DQS), provides a single summary statistic that no proposer can inflate by performing well at a cherry-picked budget. As a case study, we apply BSDS/DQS to: do LLMs add marginal value to an existing ML pipeline for drug discovery candidate selection? We evaluate 39 proposers -- 11 mechanistic variants, 14 zero-shot LLM configurations, and 14 few-shot LLM configurations -- using SMILES representations on MoleculeNet HIV (41,127 compounds, 3.5% active, 1,000 bootstrap replicates) under both random and scaffold splits. Three findings emerge. First, the simple RF-based Greedy-ML proposer achieves the best DQS (-0.046), outperforming all MLP variants and LLM configurations. Second, no LLM surpasses the Greedy-ML baseline under zero-shot or few-shot evaluation on HIV or Tox21, establishing that LLMs provide no marginal value over an existing trained classifier. Third, the proposer hierarchy generalizes across five MoleculeNet benchmarks spanning 0.18%-46.2% prevalence, a non-drug AV safety domain, and a 9x7 grid of penalty parameters (tau >= 0.636, mean tau = 0.863). The framework applies to any setting where candidates are selected under budget constraints and asymmetric error costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。