arXiv:2605.25267cs.LGcs.AI2026-05

用隐空间屏障提升安全上下文强化学习的部署表现

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

  • 预训练阶段构建隐空间动态与成本预测器
  • 部署时根据剩余预算过滤或重加权动作,提升安全与回报平衡
  • 无需更新参数,在5个基准上4项收益更高

安全上下文强化学习(ICRL)在不更新测试期参数的前提下,基于交互历史在线调整策略,同时控制每轮成本在安全预算内。当部署环境发生分布外(OOD)变化时,仅依赖预训练的安全ICRL会因冻结策略仅通过上下文条件影响行为,缺乏显式动作级成本检查,导致奖励-安全权衡不佳。本文提出一种隐空间Q-屏障盾牌,预先学习上下文表示、隐空间动态模型及集成成本判别器。部署时,该盾牌从历史中推断上下文,结合剩余预算与预测未来成本,对候选动作进行过滤或软性重加权。理论证明:满足Q-屏障的动作能使下一隐空间预算状态在学习到的成本判别器下,近似保持预算安全延续,误差由贝尔曼误差与隐空间预测误差构成。在五个安全ICRL基准上,该方法优于强基线:经过短暂上下文窗口后,四组任务获得更高回报,且全部五组平均每轮成本匹配或降低。

原文摘要 · Abstract (English)

Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under out-of-distribution (OOD) deployment shifts, pretraining-only safe ICRL can give poor reward-safety tradeoffs because the remaining budget affects behavior only through frozen policy conditioning, not an explicit action-level check against predicted future cost. We propose a latent Q-Barrier shield that learns a context representation, latent dynamics, and an ensemble cost critic before deployment. Without parameter updates, the shield infers context from history and filters or softly reweights candidate actions using the remaining budget and predicted future cost. We prove a conditional, error-decomposed barrier-margin result: a Q-Barrier-satisfying action leaves the next latent-budget state with an approximately budget-safe continuation under the learned critic, up to Bellman and latent-prediction errors. Across five safe ICRL benchmarks, the shield improves deployment-time reward-safety tradeoffs over a strong safe-ICRL baseline: after a short context window, it achieves higher return in four of five benchmarks while matching or lowering average episode cost in all five.

强化学习安全控制上下文学习隐空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。