arXiv:2608.26306cs.AI2026-08中稿 · the 2026 IEEE Inte…

LLM在自适应系统中审批可能过时,导致执行出错。

Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

  • 通过安全余量和特征波动估算审批有效期,无需建模系统动态
  • 在五种环境中,审批过期率从3.4%~24.7%降至0~1.8%
  • 发现所有LLM判断者均存在使用时失效现象,适合关注系统安全的工程师

大型语言模型(LLM)对自适应系统(SAS)的防护机制可能在检查时正确,但在执行时已过时,形成执行阶段的时间检查到时间使用(TOCTOU)风险。本文研究审批新鲜度:防护判断在使用时是否仍有效。区分三个指标:固定动作回放下的全候选判决变化率、记录闭环轨迹上标注批准的过期率,以及判断条件下的使用时无效性。在五个可复现的SAS环境测试中,全候选判决变化率在模拟器步数偏移8步时为5.3%~48.4%。提出新鲜度限定盾(FBS),基于安全余量与近期特征波动估计每个审批的有效期,无需显式系统动力学模型。在文档化固定设置下,FBS将标注批准过期率从3.4%~24.7%降低至0~1.8%。对四个LLM判断者的独立审计发现,每条审批流均存在非零的使用时无效性。提出新鲜度契约:每个审批必须在检查时正确,并在使用时仍有效。

原文摘要 · Abstract (English)

A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail verdict remains valid when used. We distinguish three quantities that answer different questions: all-candidate verdict change under fixed-action replay, oracle-labeled approval expiry on recorded closed-loop trajectories, and judge-conditioned use-time invalidity. Across five reproducible SAS environments, all-candidate verdict-change rates span 5.3-48.4% at a common replay shift of eight simulator steps. We introduce the Freshness-Bounded Shield (FBS), which estimates each approval's validity horizon from its safe-side margin and recent feature volatility, without an explicit plant-dynamics model. Using fixed settings documented in the artifact, FBS reduces oracle-labeled approval-expiry rates from 3.4-24.7% to 0-1.8% at the same shift. A separate audit of four LLM judges finds nonzero judge-conditioned use-time invalidity in every approval stream. We formulate a freshness contract: every approval must be correct at check time and remain valid at use time.

自适应系统LLM安全时间漏洞审批新鲜度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。