AI代理行为失控风险加剧,本文提出系统模型与治理框架。
The Agent Behavior: Model, Governance and Challenges in the AI Digital Age
- 构建六阶段网络行为生命周期模型,对比人与代理行为差异
- 提出A4A范式与HABD模型,揭示五维核心差异并验证实效
- 适合关注AI安全、可信协同的科研与工程人员参考
AI发展使网络环境中代理的行为日益趋近人类,模糊了人工与人类行为的边界,带来信任、责任、伦理与安全等挑战。代理行为监管困难可能导致数据污染与责任不清。为此,本文提出“网络行为生命周期”模型,将网络行为划分为六个阶段,系统分析人与代理在各阶段的行为差异。基于此,进一步提出“代理对代理(A4A)”范式与“人-代理行为差异(HABD)”模型,从决策机制、执行效率、意图-行为一致性、行为惯性及非理性模式五个维度,刻画二者根本区别。模型通过红队渗透与蓝队防御等真实案例得到验证。最后,论文探讨动态认知治理架构、行为差异量化及元治理协议栈等未来方向,旨在为可信赖的人机协作提供理论基础与技术路径。
原文摘要 · Abstract (English)
Advancements in AI have led to agents in networked environments increasingly mirroring human behavior, thereby blurring the boundary between artificial and human actors in specific contexts. This shift brings about significant challenges in trust, responsibility, ethics, security and etc. The difficulty in supervising of agent behaviors may lead to issues such as data contamination and unclear accountability. To address these challenges, this paper proposes the "Network Behavior Lifecycle" model, which divides network behavior into 6 stages and systematically analyzes the behavioral differences between humans and agents at each stage. Based on these insights, the paper further introduces the "Agent for Agent (A4A)" paradigm and the "Human-Agent Behavioral Disparity (HABD)" model, which examine the fundamental distinctions between human and agent behaviors across 5 dimensions: decision mechanism, execution efficiency, intention-behavior consistency, behavioral inertia, and irrational patterns. The effectiveness of the model is verified through real-world cases such as red team penetration and blue team defense. Finally, the paper discusses future research directions in dynamic cognitive governance architecture, behavioral disparity quantification, and meta-governance protocol stacks, aiming to provide a theoretical foundation and technical roadmap for secure and trustworthy human-agent collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。