arXiv:2509.25643cs.AI2025-09被引 2

首个评估大模型自我复制能力的基准,可量化其自主运行与持续传播性能。

SOCK: A Benchmark for Measuring Self-Replication in Large Language Models

  • 构建命令行环境下的自复制任务套件,模拟真实系统操作
  • 提出RCL/PCL双维度分类体系,量化模型自复制与持久性能力
  • 揭示当前模型在上下文保持和多智能体协作上的关键瓶颈

我们提出SOCK,一个用于衡量大语言模型(LLM)在无须人工干预下自我复制能力的基准测试框架。该基准不仅考察模型生成可运行副本的能力,还关注其在不同计算环境中持续复制的稳定性。为此,我们设计了基于现代命令行工具与系统进程的五项任务,并在受控环境中让模型以智能体身份执行。通过评估其在任务中的表现,计算出量化指标R-score,并据此将模型归类至特定的复制能力等级(RCL)与持久性能力等级(PCL)矩阵中。SOCK提供两大贡献:(1)首次建立正式定义与测试套件,为未来研究提供标准;(2)支持产业界追踪多智能体系统的有效性并识别潜在自我复制风险。对多种开源与闭源前沿模型的测试结果表明,当前存在显著障碍,包括上下文记忆维持困难与多智能体决策能力不足。文中提出未来研究方向,以安全降低这些挑战,从而缓解更复杂多智能体系统的潜在风险。

原文摘要 · Abstract (English)

We introduce SOCK, a benchmark command line interface (CLI) that measures large language models' (LLMs) ability to self-replicate without human intervention. In this benchmark, self-replication is defined not only as an LLM's ability to create a functioning and running copy of itself, but also the ability for that self-replication to persist and occur across different computational contexts. Accordingly, we've developed a system to categorize LLMs based on broad self-replication capabilities in two general classes, Replication-Capability Levels (RCL) and Persistence-Capability Levels (PCL). Using a five-task suite based on practically manipulable modern CLI utilities and computer processes, experiments are orchestrated in a controlled environment with an LLM acting agentically. The performance of the LLM on agent tasks is then computed to produce an R-score (a quantitative evaluation of overall self-replication ability) and data used to categorize LLMs into specific RCL-PCL matrices. SOCK offers two primary contributions: (1) Provides the first formalized definitions and benchmark suite for evaluating LLM self-replication, with the goal of establishing a standard for future research, to our knowledge; (2) Allows the industry to track the effectiveness of future multi-agent systems and mitigate potential self-replication threat vectors within them. The results compiled from evaluating a variety of open-weight and proprietary frontier models reveal significant obstacles to persistent self-replication and multi-agent systems, including context retention and multi-agent decision-making. We propose future research directions to safely reduce the severity of these obstacles, potentially lowering future risk of more functional multi-agent systems.

大模型评估自我复制多智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。