arXiv:2608.22512cs.AI2026-08

为自治多智能体系统设计可追责的全周期架构,防篡改、可溯源、能识别责任共谋。

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

  • 构建运行时见证机制,在五个关键节点捕获不可篡改日志
  • 通过因果图与补偿集大小量化责任归属,防责任清洗
  • 支持事故回放与责任分层报告,适合高风险自主系统

当前自治多智能体系统已在金融、软件供应链和安全运营中应用,已有部分由AI主导的入侵事件被报道。但一旦系统造成损害,现有方法无法可靠确定发生了什么、原因何在或责任归属。根源在于:溯源取证处于错误抽象层级,形式化因果依赖模型假设,代理审计信任自我记录。其典型失败模式是责任清洗——将行为分散至冗余代理,使每个个体都不构成必要条件。更严重的是,记录由嫌疑人生成,本文即基于此假设。代理可能预判调查,日志基础设施甚至可能合谋。为此提出HANSARD参考架构,将问责作为全生命周期属性。首先,运行前密封的就绪配置界定后续发现的上限;其次,在代理无法触及的五处关键点捕获数据,使遗漏可被检测,而不仅是篡改;第三,运行中累积符合PROV-DM标准的类型化因果图,实时读取三个指标以触发监督,无需裁定;第四,事故后回放基于修正的Halpern-Pearl定义,计算条件效应及补偿集规模;最后,协同残差衡量整体危害,暴露责任清洗。因果、责任与问责分别报告,各受证据层级限制,并给出未来研究方向。

原文摘要 · Abstract (English)

Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orchestrated intrusion campaigns have been reported. Yet, when such a system causes harm, no method can robustly establish what happened, what caused it, or who is accountable. This is because provenance forensics works at the wrong abstraction, formal causality assumes the causal model, and agent auditing trusts self-recording. The target failure mode is, thus, attribution laundering, i.e., spreading an act across redundant agents until none is a but-for cause. Worse, the record is produced by the suspects, which comprises the assumption adopted throughout this work. Agents may therefore anticipate the investigation and the part of logging infrastructure may itself collude. In this paper, HANSARD is proposed, a reference architecture treating accountability as a life-cycle property. First, a readiness profile sealed before operation bounds what later findings may claim. Second, capturing at five choke points beyond the agents' reach makes omissions detectable, not only tampering. Third, a typed PROV-DM-aligned causal graph accrues as the system runs, and three indicators read it live to gate oversight without adjudicating. Fourth, post-incident replay yields contingent effects under the modified Halpern-Pearl definition, together with a compensation-set size. Finally, a synergy residual measures harm due to the combination rather than to individuals, making laundering visible. Cause, responsibility and accountability are then reported separately, each capped by an evidentiary tier, while a future research agenda is also provided.

多智能体系统问责安全架构溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。