通过强化学习动态优化提示,实现对话模型推理时的安全自适应控制。
SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

- 用强化学习在推理阶段动态调整提示,实现安全行为调控。
- 多模型测试中显著提升安全性与回复质量,优于现有提示优化方法。
- 无需重训练或修改参数,适合实际部署场景中的安全管控。
确保大语言模型在真实应用中的安全与情境适切性仍是关键挑战。本文提出SafeCtrl-RL,一种推理时行为控制框架,可在不重新训练或修改模型参数的情况下实现自适应安全调节。该方法将对话生成建模为序列决策过程,由强化学习代理根据上下文反馈动态选择提示调整策略,通过迭代优化抑制不当行为,我们将其概念化为推理时的行为去学习。在多个大语言模型和不安全对话场景下的评估显示,SafeCtrl-RL持续提升安全性和响应质量,优于现有提示优化方法,并实现良好的性能-效率权衡。注意:本文可能包含有害语言内容,建议谨慎阅读。
原文摘要 · Abstract (English)
Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We present \textbf{SafeCtrl-RL}, an inference-time behavioural control framework that enables adaptive safety regulation without model retraining or parameter modification. The method formulates dialogue generation as a sequential decision process, where a reinforcement learning agent dynamically selects prompt adjustment strategies based on contextual feedback. This allows unsafe behaviours to be suppressed through iterative refinement, which we conceptualise as inference-time behavioural unlearning. Evaluated across multiple LLMs and unsafe dialogue scenarios, SafeCtrl-RL consistently improves safety and response quality, outperforms existing prompt-based optimisation methods, and achieves favourable performance--efficiency trade-offs. **Warning: This paper may contain examples of harmful language, and reader discretion is recommended.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。