arXiv:2601.01816cs.AI2026-01

用概率分布评估决策,让AI在不确定中做出更可信的选择。

Admissibility Alignment

  • 将对齐定义为基于结果分布的可接受性选择,而非静态约束。
  • 通过蒙特卡洛估算未来可能性,评估政策的期望效用与风险水平。
  • 适合企业级AI治理,尤其关注长期风险与不确定性下的决策可靠性。

本文提出Admissibility Alignment:将AI对齐重新定义为在不确定性下对可接受行动与决策选择的属性,通过候选策略的行为进行评估。我们提出MAP-AI(Monte Carlo Alignment for Policy)作为实现该对齐的典型系统架构,将对齐形式化为概率性、决策论属性,而非静态或二元条件。MAP-AI是一种新的控制平面架构,通过蒙特卡洛估计结果分布和可接受性控制的策略选择,而非静态模型约束来实现对齐。该框架在多种可能未来中评估决策策略,显式建模不确定性、干预效应、价值模糊性和治理约束。对齐通过分布特性如期望效用、方差、尾部风险及误对齐概率进行评估,而非准确率或排序性能。此方法区分了概率预测与不确定性下的决策推理,并为评估企业与机构级AI系统的信任与对齐提供了可执行的方法。最终结果是,一种实践基础,其决定影响的不是单一预测,而是策略在分布和尾部事件中的行为表现。最后,我们展示如何将分布对齐评估融入决策过程本身,实现无需重训练或修改底层模型的可接受性控制动作选择机制。

原文摘要 · Abstract (English)

This paper introduces Admissibility Alignment: a reframing of AI alignment as a property of admissible action and decision selection over distributions of outcomes under uncertainty, evaluated through the behavior of candidate policies. We present MAP-AI (Monte Carlo Alignment for Policy) as a canonical system architecture for operationalizing admissibility alignment, formalizing alignment as a probabilistic, decision-theoretic property rather than a static or binary condition. MAP-AI, a new control-plane system architecture for aligned decision-making under uncertainty, enforces alignment through Monte Carlo estimation of outcome distributions and admissibility-controlled policy selection rather than static model-level constraints. The framework evaluates decision policies across ensembles of plausible futures, explicitly modeling uncertainty, intervention effects, value ambiguity, and governance constraints. Alignment is assessed through distributional properties including expected utility, variance, tail risk, and probability of misalignment rather than accuracy or ranking performance. This approach distinguishes probabilistic prediction from decision reasoning under uncertainty and provides an executable methodology for evaluating trust and alignment in enterprise and institutional AI systems. The result is a practical foundation for governing AI systems whose impact is determined not by individual forecasts, but by policy behavior across distributions and tail events. Finally, we show how distributional alignment evaluation can be integrated into decision-making itself, yielding an admissibility-controlled action selection mechanism that alters policy behavior under uncertainty without retraining or modifying underlying models.

AI对齐决策理论不确定性风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。