arXiv:2602.23329cs.AIcs.CL2026-02被引 4

LLM让新手在生物任务上表现远超仅用网络的对照组,接近甚至超越专家。

LLM Novice Uplift on Dual-Use, In Silico Biology Tasks

  • 对比新手使用LLM与仅用互联网,评估其在生物任务上的提升效果。
  • 使用LLM的新手准确率是对照组的4.16倍,3/4任务超越专家。
  • 多数人轻松获取双用途信息,凸显安全风险与持续评估必要性。

大型语言模型(LLMs)在生物学基准测试中表现日益出色,但尚不清楚它们是否能真正提升新手能力——即帮助人类在不依赖专业训练的情况下,表现优于仅使用互联网资源者。这一问题对科学加速与双用途风险理解至关重要。我们开展了一项多模型、多基准的人类提升研究,比较了新手在拥有LLM与仅使用互联网两种条件下,在八个与生物安全相关的任务集上的表现。参与者有充足时间完成复杂任务(最长达13小时)。结果显示,使用LLM的新手准确率比对照组高出4.16倍(95%置信区间[2.63, 6.87])。在四个具备专家基准(仅互联网)的任务中,使用LLM的新手在其中三个超越了专家。出人意料的是,独立运行的LLM常优于人机协作版本,表明用户未能充分调动模型潜力。89.6%的参与者表示尽管有防护措施,仍轻松获取了双用途相关情报。总体而言,LLM显著提升了新手在原本仅限专业人士完成的生物学任务上的表现,强调必须持续进行互动式提升评估,而不仅依赖传统基准。

原文摘要 · Abstract (English)

Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurity-relevant task sets. Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). We found that LLM access provided substantial uplift: novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Perhaps surprisingly, standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Overall, LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners, underscoring the need for sustained, interactive uplift evaluations alongside traditional benchmarks.

大模型生物信息双用途风险人类增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。