用阈值法实时监控大模型输出安全,效果媲美复杂方法
Online Safety Monitoring for LLMs

- 通过外部验证器信号加阈值判断输出是否安全
- 在数学推理和红队测试数据集上表现优于传统方法
- 适合需要快速部署的安全监控场景
尽管经过对齐训练,大模型在部署时仍可能生成不安全内容。因此,在线实时监控输出并触发警报至关重要。本文研究一种简单高效的实时监控机制:将外部模型的验证信号通过阈值化处理转化为警报决策,并利用风险控制方法校准阈值。在数学推理和红队测试数据集上的实验表明,该方法性能与基于序列假设检验的先进监控方案相当。
原文摘要 · Abstract (English)
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。