arXiv:2606.13884cs.AI2026-06被引 4

用因果风险控制让大模型决策更安全,避免高成本错误

Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents

论文配图:Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents
图 1 · 摘自论文原文
  • 基于因果效应估计决定是否执行、推迟或放弃模型预测
  • 在真实和模拟场景中,高成本错误减少且性能损失小
  • 适合对安全性要求高的自动化系统,如医疗、金融决策

现代决策系统越来越多依赖于可能自信但错误的模型输出,导致下游动作产生高昂错误。本文提出风险感知因果门控(RACG),通过结合因果效应估计与校准的风险控制,决定是否执行、延迟或拒绝模型预测。RACG 建模从候选动作到结果的因果路径,并根据估算的反事实风险而非原始置信度进行决策门控。为确保门控可靠性,我们推导出在高风险条件下采取行动的概率的无分布界,并将其转化为满足用户指定安全约束的操作阈值。此外,我们提出一种自适应门控策略,通过监测预测与实际结果之间的偏差,在因果假设失效时收紧门控。在模拟干预和真实世界决策基准上,RACG 显著降低高成本错误,同时保留大部分无门控策略的效用,且在相同拒答率下优于基于置信度和选择性预测的基线方法。结果表明,明确区分因果风险与预测不确定性,可构建更安全、更透明的决策系统,为高风险场景中的可信自动化提供原则性机制。

原文摘要 · Abstract (English)

Modern decision systems increasingly rely on learned components whose outputs may be confident yet wrong, exposing downstream actions to costly errors. We introduce Risk-Aware Causal Gating (RACG), a framework that decides whether to act on, defer, or abstain from a model's prediction by combining causal effect estimation with calibrated risk control. RACG models the causal pathway from candidate actions to outcomes and gates each decision according to an estimated counterfactual risk rather than raw predictive confidence. To make gating reliable, we derive distribution-free bounds on the probability of acting under high-risk conditions and show how these bounds translate into operating thresholds that satisfy user-specified safety constraints. We further propose an adaptive gating policy that adjusts to distribution shift by monitoring discrepancies between predicted and realized outcomes, tightening the gate when causal assumptions appear violated. Across simulated interventions and real-world decision benchmarks, RACG reduces high-cost errors substantially while preserving most of the utility of an ungated policy, and it outperforms confidence-based and selective-prediction baselines at matched abstention rates. Our results indicate that explicitly separating causal risk from predictive uncertainty yields decision systems that are both safer and more transparent, offering a principled mechanism for trustworthy automation in high-stakes settings.

大模型安全因果推理决策系统风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。