arXiv:2602.01942cs.CRcs.AI2026-02被引 2

为开放环境中的智能体安全设计社会启发的四维框架。

Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework

  • 借鉴人类社会治理,构建核心、连接、认知、合规四维度安全框架
  • 覆盖智能体自主规划、协作与持续行为带来的新型风险
  • 适合研究可信智能体系统或AI治理的开发者与政策制定者

AI正从封闭可预测环境中的专用自主系统,演变为由大语言模型驱动、在开放跨组织环境中规划与行动的智能体。这类系统具备长期规划、协作与持续行为能力,成为复杂人机技术生态系统的参与者。尽管现有工作强化了对提示注入、数据污染和工具误用等模型与流程层面漏洞的防御,但以系统为中心的方法难以捕捉由自主性、交互性和涌现行为引发的风险。本文提出受人类社会治理启发的4C框架,从四个相互关联的维度系统化识别智能体风险:核心(系统、基础设施与环境完整性)、连接(通信、协调与信任)、认知(信念、目标与推理完整性)以及合规(伦理、法律与制度治理)。该框架将安全重心从单一系统保护转向行为完整性和意图一致性维护,补充现有策略,为构建可信、可控且符合人类价值的智能体系统提供原则性基础。

原文摘要 · Abstract (English)

AI is moving from domain-specific autonomy in closed, predictable settings to large-language-model-driven agents that plan and act in open, cross-organizational environments. As a result, the cybersecurity risk landscape is changing in fundamental ways. Agentic AI systems can plan, act, collaborate, and persist over time, functioning as participants in complex socio-technical ecosystems rather than as isolated software components. Although recent work has strengthened defenses against model and pipeline level vulnerabilities such as prompt injection, data poisoning, and tool misuse, these system centric approaches may fail to capture risks that arise from autonomy, interaction, and emergent behavior. This article introduces the 4C Framework for multi-agent AI security, inspired by societal governance. It organizes agentic risks across four interdependent dimensions: Core (system, infrastructure, and environmental integrity), Connection (communication, coordination, and trust), Cognition (belief, goal, and reasoning integrity), and Compliance (ethical, legal, and institutional governance). By shifting AI security from a narrow focus on system-centric protection to the broader preservation of behavioral integrity and intent, the framework complements existing AI security strategies and offers a principled foundation for building agentic AI systems that are trustworthy, governable, and aligned with human values.

智能体安全4C框架人工智能治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。