当人机协作系统自主性超过阈值,责任归属将无法完全分配。
The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

- 构建人机集体的因果模型,用四维信息论衡量自主性。
- 证明超过责任边界时,四大问责原则无法同时满足。
- 实验验证3000个合成系统,结果零偏差,揭示治理临界点。
现有人工智能系统的问责框架依赖于一个共同假设:对于任何重要后果,至少存在一个具备足够参与度与预见性的可识别个体应承担实质性责任。本文证明,一旦智能体的自主性超过一个可计算的阈值,该假设将被智能代理系统从数学上彻底打破,而非工程缺陷。我们提出「人机集体」概念,将其形式化为共享结构因果模型下的状态-策略对组合,以四维信息论指标(认知、执行、评估、社会)刻画自主性,并通过交互图与联合行动空间描述集体行为。通过四个最小化问责属性——可归因性(责任需有因果贡献)、预见性边界(责任不能超出预测能力)、非空性(至少一人负非平凡责任)、完备性(所有责任须完全分配)——建立合法问责的公理体系。核心结论:问责不完备定理指出,当集体复合自主性超过问责边界,且交互图中存在人机反馈环时,任何框架都无法同时满足四项属性。该不可能性是结构性的,透明度、审计与监督无法解决,除非降低自主性。低于阈值时,合法框架存在,形成清晰相变。3000个合成集体实验验证全部预测,无一例外。这是首个AI治理中的不可能性定理,确立了当前范式有效的下限,以及之上必须采用分布式问责机制的分界线。
原文摘要 · Abstract (English)
Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at least one identifiable person had enough involvement and foresight to bear meaningful responsibility. This paper proves that agentic AI systems violate this assumption not as an engineering limitation but as a mathematical necessity once autonomy exceeds a computable threshold. We introduce Human-Agent Collectives, a formalisation of joint human-AI systems where agents are modelled as state-policy tuples within a shared structural causal model. Autonomy is characterised through a four-dimensional information-theoretic profile (epistemic, executive, evaluative, social); collective behaviour through interaction graphs and joint action spaces. We axiomatise legitimate accountability through four minimal properties: Attributability (responsibility requires causal contribution), Foreseeability Bound (responsibility cannot exceed predictive capacity), Non-Vacuity (at least one agent bears non-trivial responsibility), and Completeness (all responsibility must be fully allocated). Our central result, the Accountability Incompleteness Theorem, proves that for any collective whose compound autonomy exceeds the Accountability Horizon and whose interaction graph contains a human-AI feedback cycle, no framework can satisfy all four properties simultaneously. The impossibility is structural: transparency, audits, and oversight cannot resolve it without reducing autonomy. Below the threshold, legitimate frameworks exist, establishing a sharp phase transition. Experiments on 3,000 synthetic collectives confirm all predictions with zero violations. This is the first impossibility result in AI governance, establishing a formal boundary below which current paradigms remain valid and above which distributed accountability mechanisms become necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。