arXiv:2602.14606cs.MAcs.AI2026-02被引 2

给智能体设计可审计的决策选择权边界,防止隐性操控。

Towards Selection as Power: Bounding Decision Authority in Autonomous Agents

  • 将认知、选择、行动分域管理,用机械规则约束选择权。
  • 在金融场景下抵御操纵攻击,降低选择集中度,提升失败可见性。
  • 适合高风险领域,如金融监管、医疗决策等需严控自主性的场景。

自主智能体正被部署于监管严格、高风险领域,其决策可能不可逆且受制度约束。现有安全方法强调对齐、可解释性或动作过滤,但这些机制不足以管控‘选择权力’——即决定哪些选项被生成、呈现和框定的能力。本文提出一种治理架构,将认知、选择与行动分离开来,将自主性建模为主权向量。认知自主性保持自由,而选择与行动自主性通过外部机制强制的原语进行限制,且不纳入代理优化空间。该架构整合了外部候选生成(CEFL)、受控缩减器、提交-揭示熵隔离、理由验证及故障告警熔断器。在多个金融场景下评估系统,针对方差操纵、阈值博弈、框架偏移、排序效应和熵探测等对抗性压力进行测试。指标包括选择集中度、叙事多样性、治理激活成本与故障可见性。结果表明,机械式选择治理可实现、可审计,能防止确定性结果捕获,同时保留推理能力。尽管存在概率性集中,该架构仍显著限制了选择权限,优于传统标量流程。本工作将治理重新定义为受控的因果权力,而非内部意图对齐,为无法容忍沉默失效的自主系统部署提供基础。

原文摘要 · Abstract (English)

Autonomous agentic systems are increasingly deployed in regulated, high-stakes domains where decisions may be irreversible and institutionally constrained. Existing safety approaches emphasize alignment, interpretability, or action-level filtering. We argue that these mechanisms are necessary but insufficient because they do not directly govern selection power: the authority to determine which options are generated, surfaced, and framed for decision. We propose a governance architecture that separates cognition, selection, and action into distinct domains and models autonomy as a vector of sovereignty. Cognitive autonomy remains unconstrained, while selection and action autonomy are bounded through mechanically enforced primitives operating outside the agent's optimization space. The architecture integrates external candidate generation (CEFL), a governed reducer, commit-reveal entropy isolation, rationale validation, and fail-loud circuit breakers. We evaluate the system across multiple regulated financial scenarios under adversarial stress targeting variance manipulation, threshold gaming, framing skew, ordering effects, and entropy probing. Metrics quantify selection concentration, narrative diversity, governance activation cost, and failure visibility. Results show that mechanical selection governance is implementable, auditable, and prevents deterministic outcome capture while preserving reasoning capacity. Although probabilistic concentration remains, the architecture measurably bounds selection authority relative to conventional scalar pipelines. This work reframes governance as bounded causal power rather than internal intent alignment, offering a foundation for deploying autonomous agents where silent failure is unacceptable.

自主代理治理架构金融风控决策安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。