测试大模型解决病毒实验难题的能力,发现部分模型已超专家水平。
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
- 构建322道多模态病毒实验题,涵盖基础、隐性与视觉知识
- 顶尖模型o3准确率达43.8%,超过94%专家在专长领域表现
- 模型能力具双刃剑效应,需纳入生命科学双重用途技术治理
我们提出病毒学能力测试(Virology Capabilities Test, VCT),一个大型语言模型(LLM)基准,用于评估模型在解决复杂病毒学实验室协议问题上的能力。该基准由数十位博士级病毒学家提供输入,包含322道多模态问题,覆盖病毒实验中必需的基础性、隐性及视觉知识。VCT难度极高:即使有网络支持的专家,在其专业领域平均仅答对22.1%。然而,表现最优的LLM——OpenAI的o3模型,准确率达到43.8%,超越94%的专家在各自子领域的表现。能够提供专家级病毒学故障排除能力具有双重用途:既可用于有益科研,也可能被滥用。因此,公开模型在VCT上表现优于人类专家,引发紧迫的治理挑战。我们建议将大模型在双重用途病毒学工作中的故障排除能力,纳入现有生命科学双重用途技术管理框架。
原文摘要 · Abstract (English)
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of $322$ multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of $22.1\%$ on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches $43.8\%$ accuracy, outperforming $94\%$ of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。