为大规模多智能体系统设计可追踪、可干预的责任机制,防止合谋等异常行为。
Adaptive Accountability in Networked MAS: Tracing and Mitigating Emergent Norms at Scale
- 通过因果图追踪交互责任,实时记录可验证的行为证据。
- 在96%场景中降低违规率11.9%,且社会福利几乎不受影响。
- 适合需要高可靠性的分布式系统,如自动驾驶或金融交易网络。
大规模网络化多智能体系统日益支撑关键基础设施,但其集体行为可能演变为合谋、资源囤积和隐性不公平等不良涌现规范。本文提出自适应问责框架(AAF),一个端到端的运行时层,具备:(i) 可密码验证的交互溯源记录;(ii) 流式日志中的分布变化点检测;(iii) 基于因果影响图的责任归因;(iv) 成本受限的干预措施——奖励重塑与定向策略修补,以引导系统回归合规行为。我们证明了有界妥协保证:若干预预期成本超过攻击者预期收益,则长期违规交互比例收敛至低于1的值。在包含87,480次运行的大规模因子仿真中(两任务;最多100个智能体+500个扩展测试;全/部分可观测;拜占庭率最高10%;每种情形10个随机种子),在324种配置下,与近端策略优化基线相比,AAF在96%场景中降低了执行违规率(中位相对降幅11.9%),同时保持社会福利稳定(中位变化0.4%)。面对恶意注入,AAF在中位71步内检测到规范违规(四分位距39–177),在10%拜占庭率下实现0.97的平均最高排名归因准确率。
原文摘要 · Abstract (English)
Large-scale networked multi-agent systems increasingly underpin critical infrastructure, yet their collective behavior can drift toward undesirable emergent norms such as collusion, resource hoarding, and implicit unfairness. We present the Adaptive Accountability Framework (AAF), an end-to-end runtime layer that (i) records cryptographically verifiable interaction provenance, (ii) detects distributional change points in streaming traces, (iii) attributes responsibility via a causal influence graph, and (iv) applies cost-bounded interventions-reward shaping and targeted policy patching-to steer the system back toward compliant behavior. We establish a bounded-compromise guarantee: if the expected cost of intervention exceeds an adversary's expected payoff, the long-run fraction of compromised interactions converges to a value strictly below one. We evaluate AAF in a large-scale factorial simulation suite (87,480 runs across two tasks; up to 100 agents plus a 500-agent scaling sweep; full and partial observability; Byzantine rates up to 10%; 10 seeds per regime). Across 324 regimes, AAF lowers the executed compromise ratio relative to a Proximal Policy Optimization baseline in 96% of regimes (median relative reduction 11.9%) while preserving social welfare (median change 0.4%). Under adversarial injections, AAF detects norm violations with a median delay of 71 steps (interquartile range 39-177) and achieves a mean top-ranked attribution accuracy of 0.97 at 10% Byzantine rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。