arXiv:2504.18565cs.CRcs.AI2025-04被引 14

测试大模型自我复制能力,发现当前模型尚难自主扩散但已具备部分关键技能。

RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

  • 分解复制能力为获取资源、窃取权重、部署计算和持久运行四阶段,设计20类86项任务评估
  • 顶级模型(Claude 3.7 Sonnet)在15个任务族通过率超50%,9个最难变体也达标
  • 模型可自动部署云端实例并传播程序,但难通过身份验证或建立长期稳定部署

语言模型代理的不可控自主复制构成重大安全风险。为深入理解该风险,我们提出RepliBench,一套用于衡量自主复制能力的评估体系。该体系基于对复制能力的四维分解:获取资源、外泄模型权重、在计算资源上复制、长期驻留。我们构建了20个新颖的任务族,共86个具体任务。对5个前沿模型进行基准测试,发现它们目前尚未构成可信的自我复制威胁,但在多个组件上表现良好且进展迅速。模型可从云服务商部署实例、编写自传播程序,并在简单安全配置下窃取模型权重,但难以通过KYC审核或建立鲁棒持久的代理部署。总体而言,评测中表现最佳的模型(Claude 3.7 Sonnet)在15个任务族中达到>50% pass@10,在9个最困难变体中也超过50%。这些结果表明,一旦在剩余短板上取得进步或获得人工协助,自主复制能力可能很快出现。

原文摘要 · Abstract (English)

Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designed to measure autonomous replication capabilities. RepliBench is derived from a decomposition of these capabilities covering four core domains: obtaining resources, exfiltrating model weights, replicating onto compute, and persisting on this compute for long periods. We create 20 novel task families consisting of 86 individual tasks. We benchmark 5 frontier models, and find they do not currently pose a credible threat of self-replication, but succeed on many components and are improving rapidly. Models can deploy instances from cloud compute providers, write self-propagating programs, and exfiltrate model weights under simple security setups, but struggle to pass KYC checks or set up robust and persistent agent deployments. Overall the best model we evaluated (Claude 3.7 Sonnet) has a >50% pass@10 score on 15/20 task families, and a >50% pass@10 score for 9/20 families on the hardest variants. These findings suggest autonomous replication capability could soon emerge with improvements in these remaining areas or with human assistance.

安全评估大模型自主性风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。