arXiv:2609.06445cs.AI2026-09

为智能体决策提供可追溯的因果归因框架,解决高风险AI系统审计难题。

Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification

  • 提出分离边际与共随机数效应的估计框架,区分步骤真实贡献。
  • 发现关键步骤在特定情况下被误判为无影响,暴露现有方法根本缺陷。
  • 构建耦合机制保障直接效应可估,适用于需严格可解释性的系统开发。

高风险AI系统提供方须保留可追溯决策记录,但针对智能体系统尚未明确记录应包含何种内容以支持事后因果归因。本文提出估计器框架,并分析其失效条件。将先前工作测量的边际总效应与基于共随机数的总效应相区分,引入下游固定时的自然直接效应,并通过人工推导验证估计器表现。两者均失败,且方向一致:在设定链中,因果无关步骤的总效应与决定性步骤完全相同,这是代数恒等而非偶然;在共随机数下,决定性步骤在执行步骤翻转的约1/10情形中总效应精确为零,而其直接效应为0.25,证明其确有作用。因此,零效应不等于无行为。本文推导出在上下文分歧后仍可估计直接效应的耦合机制,并给出其退化闭式表达。所构建的中介占比在抑制情形下超过1,导致排名异常——被抑制成分反而高于纯中介。因实验所需实时流水线未在研究期内可用,仅发布预注册协议。最后,提出符合《法案》日程空档的可追溯性规范要求:第86条解释权自2026年8月2日起生效,而第12条日志与附件四文档需至2027年12月2日才生效。

原文摘要 · Abstract (English)

A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditions under which it fails. We separate the marginal total effect that prior work measures from a common-random-numbers total effect that isolates a step's own contribution, add the natural direct effect under a pinned downstream, and check the estimators against hand derivations. Both estimands then fail, in the same direction. Under the marginal estimand a causally inert step has the identical total effect to the decisive one on every run of our planted chain, an algebraic identity and not a coincidence at one draw. Under common random numbers the decisive step returns exactly zero on the runs where the executing step flips, about one in ten, while its direct effect there is 0.25 and it demonstrably acts; an exact zero does not certify that a step did nothing, and we put that here rather than in the limitations. We derive the coupling that keeps the direct effect estimable once contexts diverge, with a closed form for its degradation, and show that the mediated share on which a natural ranking is built is not a share under suppression: where the direct and mediated paths oppose, it exceeds one and ranks a suppressed component above a pure mediator. We publish the discrepancy experiment's pre-registration rather than a result, because the live pipeline it requires was not available in the study window. We contribute the traceability specification such a filing would need, against a gap the Act's calendar opens: Article 86's right to an explanation has applied since 2 August 2026, while the Article 12 logging and Annex IV documentation that could evidence one were deferred to 2 December 2027 by Regulation (EU) 2026/1744.

因果归因智能体系统可解释性合规审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。