arXiv:2505.06108cs.LGcs.AI2025-05被引 9

大模型在生物领域超越专家,尤其在病毒学测试中表现卓越。

LLMs Outperform Experts on Challenging Biology Benchmarks

  • 用8个生物基准测试27个前沿大模型,每项测10次确保可靠性。
  • 顶级模型在病毒学文本测试中表现提升超4倍,o3模型能力达专家两倍。
  • 部分模型已接近或超过专家水平,但现有评测存在数据错误和饱和问题。

本研究系统评估了27个前沿大语言模型在涵盖分子生物学、遗传学、克隆、病毒学和生物安全的8个生物基准上的表现,模型来自主要AI开发者,发布时间介于2022年11月至2025年4月之间,每个基准进行十次独立运行。结果表明生物能力显著提升:在病毒学能力测试的纯文本子集上,顶尖模型性能在研究周期内增长超过4倍,OpenAI的o3模型表现已达专家病毒学家的两倍。多个模型在GPQA和WMDP的生物子集及LAB-Bench克隆场景等挑战性任务中达到或超过专家水平。出乎意料的是,思维链提示并未显著优于零样本评估;而o3-mini和Claude 3.7 Sonnet的扩展推理功能则按预期随推理规模提升。然而,PubMedQA、MMLU和WMDP生物子集表现出性能平台期,远低于100%,暗示基准饱和与数据本身存在错误。研究强调需发展更先进的评估方法以应对持续演进的AI系统。

原文摘要 · Abstract (English)

This study systematically evaluates 27 frontier Large Language Models on eight biology benchmarks spanning molecular biology, genetics, cloning, virology, and biosecurity. Models from major AI developers released between November 2022 and April 2025 were assessed through ten independent runs per benchmark. The findings reveal dramatic improvements in biological capabilities. Top model performance increased more than 4-fold on the challenging text-only subset of the Virology Capabilities Test over the study period, with OpenAI's o3 now performing twice as well as expert virologists. Several models now match or exceed expert-level performance on other challenging benchmarks, including the biology subsets of GPQA and WMDP and LAB-Bench CloningScenarios. Contrary to expectations, chain-of-thought did not substantially improve performance over zero-shot evaluation, while extended reasoning features in o3-mini and Claude 3.7 Sonnet typically improved performance as predicted by inference scaling. Benchmarks such as PubMedQA and the MMLU and WMDP biology subsets exhibited performance plateaus well below 100%, suggesting benchmark saturation and errors in the underlying benchmark data. The analysis highlights the need for more sophisticated evaluation methodologies as AI systems continue to advance.

大模型生物医学评估基准AI专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。