arXiv:2503.15182cs.CYcs.AI2025-03被引 1

测试发现大模型对新型生化威胁的披露能力分阶段进化,专家引导下表现差异显著。

Foundation models may exhibit staged progression in novel CBRN threat disclosure

  • 通过模拟新生物威胁披露,对比大模型与搜索结果的推理表现
  • 高级模型在专家提示下正确推理,低级模型即使引导也失败(80 vs 5)
  • 揭示模型能力有四个阶段,适合安全研究人员关注其进展

由于缺乏测试案例,基础模型向专家用户披露新型化学、生物、辐射和核(CBRN)威胁的能力尚不明确。本文利用即将公开的《镜像细菌:可行性与风险》技术报告,开展了一项小型受控研究。训练过的生物学家使用Claude Sonnet 3.5(n=10)或仅用网络搜索(n=2)预测释放镜像大肠杆菌的后果,两者评分无显著差异(分别为28和43),均低于网络基线(36)。但当由报告作者提示时,Sonnet能正确推理,而较小的Haiku 3.5模型即便在引导下仍失败(80 vs 5)。结果表明模型能力存在阶段性:Haiku无法在专家引导下推理镜像生命(第1阶段),Sonnet仅在威胁敏感提示下可正确推理(第2阶段)。未来模型可能逐步实现向普通专家(第3阶段)或非专业人士(第4阶段)披露新威胁。尽管镜像生命仅为个案,持续监测模型对私有威胁的推理能力,有助于在广泛披露前采取防护措施。

原文摘要 · Abstract (English)

The extent to which foundation models can disclose novel chemical, biological, radiation, and nuclear (CBRN) threats to expert users is unclear due to a lack of test cases. I leveraged the unique opportunity presented by an upcoming publication describing a novel catastrophic biothreat - "Technical Report on Mirror Bacteria: Feasibility and Risks" - to conduct a small controlled study before it became public. Graduate-trained biologists tasked with predicting the consequences of releasing mirror E. coli showed no significant differences in rubric-graded accuracy using Claude Sonnet 3.5 new (n=10) or web search only (n=2); both groups scored comparably to a web baseline (28 and 43 versus 36). However, Sonnet reasoned correctly when prompted by a report author, but a smaller model, Haiku 3.5, failed even with author guidance (80 versus 5). These results suggest distinct stages of model capability: Haiku is unable to reason about mirror life even with threat-aware expert guidance (Stage 1), while Sonnet correctly reasons only with threat-aware prompting (Stage 2). Continued advances may allow future models to disclose novel CBRN threats to naive experts (Stage 3) or unskilled users (Stage 4). While mirror life represents only one case study, monitoring new models' ability to reason about privately known threats may allow protective measures to be implemented before widespread disclosure.

大模型安全生物威胁推理能力风险披露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。