用可验证工作流筛选生物新发现,确保结果可追溯、可审计。
Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

- 构建带溯源记录的生物学专用智能体,链式追踪证据与主张关联
- 在15个人类蛋白上,11个预测结构误差小于1埃,显著优于旧模型
- 适合需严格验证的研究者或需要可复现结论的生物医学项目
大语言模型研究代理虽能整合文献检索、分析代码和论文撰写,但输出连贯不等于科学有效。我们开发了Plato-Bio,基于开源Plato/Denario架构的生物学专用扩展,通过显式工作流状态、溯源记录、引用核查、主张-证据关联、作用域文件写入及发表门控机制,保障过程可审计。源审计发现并修复三处缺陷:默认工厂中任务领域丢失、评分未纳入声明方法信号、证据附加文件缺少已声明主张的基数。经修正后,完整Python套件完成931次成功运行,6次跳过,无失败或错误;针对生物学、基因组学、证据/引用及对抗安全性的子套件亦全部成功。评估两个特定场景:在历史回溯任务中,独立前1986年文献桥接将鱼油与雷诺综合征关系排第一,TF-IDF第二,语料频率第三,仅衡量回顾性排名而非前瞻性发现。在15个人类蛋白的AlphaFold模型与实验结构对比中,11个目标的核心α碳原子均方根偏差低于1埃(中位数0.501埃),4个超过2埃;置信度掩码使SUMO1的偏差从16.61降至2.58埃(74个残基)。工作流共生成27个可追溯的差异区域,均作为待验证假设保留。因此,Plato-Bio提供可复现的软件契约和可审计的筛查基准;对代理效能或生物新发现的广泛宣称,仍需预注册评估、独立审查和前瞻性验证。
原文摘要 · Abstract (English)
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario architecture that couples explicit workflow states with provenance records, citation checks, claim-to-evidence links, scoped file writes, and publication gates. A source audit identified and repaired three defects that could distort evaluation: loss of task domain in the default factory, omission of declared method signals from scoring, and evidence sidecars that lacked the drafted-claim denominator. On the current clean revision, the full Python suite completed with 931 passes, six skips, and no failures or errors; targeted biology, genomics, evidence/citation, and adversarial-safety suites likewise completed without failure. We evaluated two narrow use cases. In a frozen historical rediscovery task, independent pre-1986 literature bridges ranked the later-studied relation between fish oil and Raynaud phenomenon first; TF-IDF ranked it second and corpus frequency third. This single curated task measures retrospective ranking, not prospective discovery. In a separate comparison of AlphaFold models with experimental structures for 15 human proteins, 11 targets had high-confidence-core C-alpha RMSD below 1 Angstrom (median 0.501 Angstrom). Four targets exceeded 2 Angstrom, and confidence masking reduced the SUMO1 discrepancy from 16.61 to 2.58 Angstrom over 74 residues. The workflow emitted 27 traceable discrepancy regions, all retained as unvalidated hypotheses. Plato-Bio therefore provides reproducible software contracts and auditable screening baselines; broader claims of agent efficacy or biological novelty require preregistered evaluation, independent review, and prospective validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。