arXiv:2605.16872cs.CYcs.AI2026-05

AI代理需有能承受后果的实体,否则责任将转嫁给人类。

Some[Body] Must Receive That Pain for Agent Accountability

  • 提出‘后果接收’概念:行为后果必须由具体实体承担以改变未来行为。
  • 指出当前LLM代理缺乏身体性,无法形成持续反馈机制。
  • 强调责任需设计在系统中,否则人类将被动承担后果。

AI代理在现实世界中的行动日益产生实际后果。然而,当前系统存在‘后果接收’问题:尽管伤害发生且责任主体可识别,但没有持续的代理能真正承受后果并据此调整行为。疼痛作为纠正性反馈信号,在惩罚理论(威慑、改造、报应、隔离)中具有基础作用,这要求有一个能承载信号的‘身体’——具备边界保护、信号累积、固化为持久更新、并驱动行为改变的载体。现有大模型代理(由权重、提示、工具、记忆、凭证等自由组合而成)均不满足这些条件。两种主流法律应对方式均失效:‘薄身份’模式下人类承担了非自身控制的行为后果,形成埃利什所称的‘道德缓冲区’;‘厚身份’模式虽构建了法律主体,却无法确保任何决策架构真正接收痛苦作为行为信号。因此,实现后果-代理耦合是社会技术基础设施问题,而非仅法律问题。在具备此类架构前,高风险AI部署应始终绑定具备实质性控制权、责任比例匹配、可约束或终止代理的人类负责人。若系统设计未让某实体主动承受后果,则必然有人被动承受。

原文摘要 · Abstract (English)

AI agents increasingly act consequentially in the real world. This creates a problem we call \emph{consequence reception}: harm occurs, the producing system is identified, yet no continuing agent receives consequences in a way that changes future behavior. Pain, understood mechanistically as a corrective feedback signal, is foundational to canonical theories of punishment -- deterrence, rehabilitation, retribution, and incapacitation all assume a continuing locus that registers the signal and updates behavior. That, in turn, requires a body for the signal to land on: a boundary whose integrity it protects, a locus where it accumulates, consolidation that converts episodic signal into durable update, and a substrate that responds by altering future action. Current LLM agents -- software-defined composites of weights, prompts, tools, memory, and credentials, freely swapped, copied, reset, and reassembled -- satisfy none of these conditions. The two prevailing legal responses therefore fail to achieve consequence reception. The thin-identity agent-principal dyad has a body but no \emph{consequence--agency coupling}: the human bears pain for behaviors beyond their control -- Elish's \emph{moral crumple zone}. The thick-identity Arbel et al.'s \emph{Algorithmic Corporation} creates legally legible entities but does not guarantee that any AI decision architecture receives pain as a behavioral signal. Achieving consequence-agency coupling is therefore a sociotechnical infrastructural problem, not only a legal one. Until such architectures exist, high-stakes AI deployment should remain tethered to accountable human principals with meaningful control, proportional liability, and authority to constrain or terminate the agent. \emph{If some body does not receive the pain by design, some body will receive it by default.}

AI伦理责任归属代理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。