从个体交互机制入手,揭示智能体系统集体风险的成因与防控路径。
Agentic Microphysics: A Manifesto for Generative AI Safety
- 聚焦智能体间局部交互结构,构建因果明确的分析框架。
- 发现群体风险源于多智能体持续互动中的累积效应,而非单个模型缺陷。
- 适合关注AI安全、多智能体系统设计的研究者与工程师。
本文提出一种面向代理型AI安全研究的方法论。随着系统具备规划、记忆、工具使用、持久身份和持续交互能力,安全问题不再仅可由孤立模型分析。群体层面的风险源于智能体之间的结构性互动,包括通信、观察和相互影响,这些过程随时间塑造集体行为。分析对象的转变导致方法论空白:仅关注单个智能体或整体结果的方案,无法识别生成集体风险的交互级机制及其可控变量。亟需一个将局部交互结构与群体动态以因果方式关联的框架,以实现解释与干预。本文引入两个核心概念:代理微观物理(Agentic microphysics)定义分析层级——在特定协议条件下,一个智能体的输出成为另一个输入的局部交互动态;生成安全(Generative safety)定义方法论——从微观条件出发,追踪现象生成与风险诱发过程,识别临界阈值并设计有效干预措施。
原文摘要 · Abstract (English)
This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safety can no longer be analysed primarily at the level of the isolated model. Population-level risks arise from structured interaction among agents, through processes of communication, observation, and mutual influence that shape collective behaviour over time. As the object of analysis shifts, a methodological gap emerges. Approaches focused either on single agents or on aggregate outcomes do not identify the interaction-level mechanisms that generate collective risks or the design variables that control them. A framework is required that links local interaction structure to population-level dynamics in a causally explicit way, allowing both explanation and intervention. We introduce two linked concepts. Agentic microphysics defines the level of analysis: local interaction dynamics where one agent's output becomes another's input under specific protocol conditions. Generative safety defines the methodology: growing phenomena and elicit risks from micro-level conditions to identify sufficient mechanisms, detect thresholds, and design effective interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。