测试大模型对新手生物实验帮助,发现效果有限但有微弱提升。
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
- 用随机对照试验评估大模型在病毒反向遗传学实验中的辅助效果。
- 大模型组完成率5.2%,互联网组6.6%,差异不显著(P=0.759)。
- 细胞培养任务成功率更高(68.8% vs 55.3%),适合需实操指导的新手。
大型语言模型(LLMs)在生物基准测试中表现优异,引发对其可能帮助新手掌握双用途实验技能的担忧。然而,这种能力是否能转化为真实实验室中的人类表现提升仍不明确。为此,我们开展了一项注册过的、研究者盲法的随机对照试验(2025年6月至8月;n=153),评估大模型是否能提升新手在模拟病毒反向遗传学流程的综合任务中的表现。主要终点为流程完成率(大模型组5.2% vs 互联网组6.6%;P=0.759),以及各任务的成功率。结果显示两者无显著差异。但大模型组在五项任务中有四项的完成率更高,尤其在细胞培养任务中(68.8% vs 55.3%;P=0.059)。事后贝叶斯建模显示,典型反向遗传学任务在大模型辅助下成功率约提升1.4倍(95%可信区间0.74–2.62)。序数回归分析表明,大模型组更可能完成中间步骤(后验概率81%-96%)。总体而言,2025年中期的大模型并未显著提高新手完成复杂实验流程的能力,但表现出轻微优势。结果揭示了数字基准与现实应用间的差距,强调必须持续验证人工智能在生物安全评估中的实际效能。
原文摘要 · Abstract (English)
Large language models (LLMs) perform strongly on biological benchmarks, raising concerns that they may help novice actors acquire dual-use laboratory skills. Yet, whether this translates to improved human performance in the physical laboratory remains unclear. To address this, we conducted a pre-registered, investigator-blinded, randomized controlled trial (June-August 2025; n = 153) evaluating whether LLMs improve novice performance in tasks that collectively model a viral reverse genetics workflow. We observed no significant difference in the primary endpoint of workflow completion (5.2% LLM vs. 6.6% Internet; P = 0.759), nor in the success rate of individual tasks. However, the LLM arm had numerically higher success rates in four of the five tasks, most notably for the cell culture task (68.8% LLM vs. 55.3% Internet; P = 0.059). Post-hoc Bayesian modeling of pooled data estimates an approximate 1.4-fold increase (95% CrI 0.74-2.62) in success for a "typical" reverse genetics task under LLM assistance. Ordinal regression modelling suggests that participants in the LLM arm were more likely to progress through intermediate steps across all tasks (posterior probability of a positive effect: 81%-96%). Overall, mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit. These results reveal a gap between in silico benchmarks and real-world utility, underscoring the need for physical-world validation of AI biosecurity assessments as model capabilities and user proficiency evolve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。