arXiv:2512.02682cs.MAcs.AI2025-12被引 7

当大模型互相对话时,单个模型安全无法保障整体系统,这篇论文提出新框架应对集体风险。

Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions

  • 从单模型安全转向系统级安全,关注多模型交互中的风险累积
  • 提出新兴系统性风险边界(ESRH)理论,解释为何局部合规也会导致全局失败
  • 设计InstitutionalAI架构,支持多智能体系统中动态监管

本文探讨为何为人类-模型交互设计的安全机制无法适用于大语言模型(LLMs)之间相互作用的场景。当前多数治理实践仍依赖于单代理安全约束,如提示工程、微调和内容审核层,这些方法仅规范单个模型行为,却忽视了多模型交互的动态特性。这些机制假设的是单一模型响应单一用户且处于稳定监管下,而研究与产业正快速转向由多个模型构成的生态系统,其中输出被递归用作后续输入。在此类系统中,即使每个模型均个体对齐,局部合规也可能累积成集体失效。本文提出从模型级安全向系统级安全的范式转变,引入新兴系统性风险边界(ESRH)框架,形式化地描述稳定性如何源于交互结构而非孤立偏差。论文贡献包括:(i) 对交互式LLM中集体风险的理论阐释,(ii) 连接微观、中观与宏观层面失效模式的分类体系,(iii) 为嵌入自适应监督而设计的InstitutionalAI架构。

原文摘要 · Abstract (English)

This paper examines why safety mechanisms designed for human-model interaction do not scale to environments where large language models (LLMs) interact with each other. Most current governance practices still rely on single-agent safety containment, prompts, fine-tuning, and moderation layers that constrain individual model behavior but leave the dynamics of multi-model interaction ungoverned. These mechanisms assume a dyadic setting: one model responding to one user under stable oversight. Yet research and industrial development are rapidly shifting toward LLM-to-LLM ecosystems, where outputs are recursively reused as inputs across chains of agents. In such systems, local compliance can aggregate into collective failure even when every model is individually aligned. We propose a conceptual transition from model-level safety to system-level safety, introducing the framework of the Emergent Systemic Risk Horizon (ESRH) to formalize how instability arises from interaction structure rather than from isolated misbehavior. The paper contributes (i) a theoretical account of collective risk in interacting LLMs, (ii) a taxonomy connecting micro, meso, and macro-level failure modes, and (iii) a design proposal for InstitutionalAI, an architecture for embedding adaptive oversight within multi-agent systems.

大模型安全多智能体系统风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。