提出机械良知框架,让智能系统在不确定中保持全局行为合规。
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence
- 用最小修正确保智能体轨迹不偏离规范区域。
- 实验证明其可避免传统控制器的全局违规,且支持多智能体协同。
- 适合研究分布式智能系统的可靠性与可解释治理。
分布式协同智能(DCI)包括边缘到边缘架构、联邦学习、迁移学习和群体系统,在不确定性下必然产生涌现风险:个体智能体局部正确决策可能组合成全局不可接受的行为轨迹。现有方法如约束优化、安全强化学习和运行时保障仅评估单个动作的可接受性,未涵盖多参与方、高不确定性的DCI场景。本文提出机械良知(MC),一种新概念与简化数学框架,实现对单智能体及分布式系统在轨迹层面的规范监管。机械良知定义为监督过滤器,以最小代价修正基线策略的动作,降低累积偏离规范区域的程度,同时考虑认知不确定性。文中引入良知得分、机械内疚与共振可靠性等概念,提供可解释的治理语言与可计算信号。理论证明了规范等价性、最优调控存在性及偏差单调递减性。实例表明,经MC调节的智能体能维持轨迹层面规范可接受性,而传统控制器会偏离可接受边界;该框架还能自然抑制多智能体环境下由交互引发的涌现风险。
原文摘要 · Abstract (English)
Distributed collaborative intelligence (DCI), encompassing edge-to-edge architectures, federated learning, transfer learning, and swarm systems, creates environments in which emergent risk is structurally unavoidable: locally correct decisions by individual agents compose into globally unacceptable behavioral trajectories under uncertainty. Existing approaches such as constrained optimization, safe reinforcement learning, and runtime assurance evaluate acceptability at the level of individual actions rather than across behavioral trajectories, and none addresses the multi-participant, uncertainty-laden nature of DCI deployments. This paper introduces mechanical conscience (MC), a novel concept and simplified mathematical framework that operationalizes trajectory-level normative regulation for both single-agent and distributed intelligent systems. Mechanical conscience is defined as a supervisory filter that minimally corrects a baseline policy's actions to reduce cumulative deviation from a normatively admissible region, while accounting for epistemic uncertainty. We introduce associated constructs, conscience score, mechanical guilt, and resonant dependability, that provide an interpretable vocabulary and computable governance signals for this emerging field. Core theoretical properties are established: admissibility equivalence, existence of optimal regulation, and monotonic deviation reduction. Illustrative results demonstrate that MC-regulated agents maintain trajectory-level normative acceptability where conventional controllers drift outside admissible bounds, and that the framework naturally extends to suppress interaction-induced emergent risk in multi-agent DCI settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。