arXiv:2510.13727cs.AI2025-10被引 3

用控制理论让AI主动避险,不再只会拒绝任务。

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails

  • 将AI安全建模为连续决策问题,通过控制理论实时纠正风险动作。
  • 在模拟驾驶和电商场景中避免碰撞与破产,任务性能不受影响。
  • 适用于任何大模型,可动态修复而非简单拒绝,适合高危应用。

生成式AI系统正越来越多地在实际场景中代表用户行动,从数字购物助手到下一代自动驾驶汽车。在此背景下,安全不再仅是阻止有害内容,而是预防下游危害,如财务或身体伤害。然而,大多数AI防护机制仍依赖基于标注数据的输出分类和人工设定标准,对新危险情境缺乏鲁棒性。即使检测到不安全状态,也无恢复路径:通常系统只是拒绝执行,而这未必安全。本文提出,智能体式AI安全本质上是一个序列决策问题,有害结果源于系统持续交互及其对世界产生的连锁后果。我们通过安全关键控制理论视角,在AI模型的潜在表征空间中形式化这一问题,构建可预测的防护机制:(i) 实时监控AI输出(行为),(ii) 主动将风险输出修正为安全输出,且方法模型无关,可适配任意AI模型。我们还提供一种基于安全关键强化学习的大规模训练方案。在模拟驾驶和电子商务场景中的实验表明,该控制理论防护机制能可靠规避灾难性后果(如碰撞、破产),同时保持任务性能,为当前‘标记并阻止’式防护提供了原则性动态替代方案。

原文摘要 · Abstract (English)

Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-generation autonomous cars. In this context, safety is no longer about blocking harmful content, but about preempting downstream hazards like financial or physical harm. Yet, most AI guardrails continue to rely on output classification based on labeled datasets and human-specified criteria,making them brittle to new hazardous situations. Even when unsafe conditions are flagged, this detection offers no path to recovery: typically, the AI system simply refuses to act--which is not always a safe choice. In this work, we argue that agentic AI safety is fundamentally a sequential decision problem: harmful outcomes arise from the AI system's continually evolving interactions and their downstream consequences on the world. We formalize this through the lens of safety-critical control theory, but within the AI model's latent representation of the world. This enables us to build predictive guardrails that (i) monitor an AI system's outputs (actions) in real time and (ii) proactively correct risky outputs to safe ones, all in a model-agnostic manner so the same guardrail can be wrapped around any AI model. We also offer a practical training recipe for computing such guardrails at scale via safety-critical reinforcement learning. Our experiments in simulated driving and e-commerce settings demonstrate that control-theoretic guardrails can reliably steer LLM agents clear of catastrophic outcomes (from collisions to bankruptcy) while preserving task performance, offering a principled dynamic alternative to today's flag-and-block guardrails.

AI安全控制理论生成式AI动态防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。