arXiv:2410.10414cs.CRcs.CL2024-10ICLR被引 25

检测大模型内容过滤器的误判问题,提升安全审核可靠性

On Calibration of LLM-based Guard Models for Reliable Content Moderation

  • 测试9个大模型过滤器在12个数据集上的置信度校准情况
  • 发现现有过滤器普遍过度自信,对抗攻击下表现严重失准
  • 提出用温度缩放和上下文校准改善置信度,尤其无验证集时有效

大语言模型可能生成有害内容或被用户绕过安全限制。已有研究开发了基于LLM的过滤模型,用于监管威胁性模型的输入输出,确保符合安全策略。然而,对这类过滤模型的可靠性与置信度校准关注不足。本文针对9个现有的基于LLM的过滤模型,在12个基准数据集上对用户输入与模型输出分类任务进行了全面的置信度校准实验。结果表明:当前过滤模型普遍存在过度自信、在越狱攻击下显著失准、且对不同响应模型输出的鲁棒性有限等问题。我们评估了后处理校准方法的效果,证明温度缩放有效,并首次揭示上下文校准在缺乏验证集时对置信度校准的益处。分析与实验凸显了现有过滤模型的局限性,为未来开发更可靠的内容过滤系统提供了重要启示。建议未来发布过滤模型时纳入置信度校准的可靠性评估。

原文摘要 · Abstract (English)

Large language models (LLMs) pose significant risks due to the potential for generating harmful content or users attempting to evade guardrails. Existing studies have developed LLM-based guard models designed to moderate the input and output of threat LLMs, ensuring adherence to safety policies by blocking content that violates these protocols upon deployment. However, limited attention has been given to the reliability and calibration of such guard models. In this work, we empirically conduct comprehensive investigations of confidence calibration for 9 existing LLM-based guard models on 12 benchmarks in both user input and model output classification. Our findings reveal that current LLM-based guard models tend to 1) produce overconfident predictions, 2) exhibit significant miscalibration when subjected to jailbreak attacks, and 3) demonstrate limited robustness to the outputs generated by different types of response models. Additionally, we assess the effectiveness of post-hoc calibration methods to mitigate miscalibration. We demonstrate the efficacy of temperature scaling and, for the first time, highlight the benefits of contextual calibration for confidence calibration of guard models, particularly in the absence of validation sets. Our analysis and experiments underscore the limitations of current LLM-based guard models and provide valuable insights for the future development of well-calibrated guard models toward more reliable content moderation. We also advocate for incorporating reliability evaluation of confidence calibration when releasing future LLM-based guard models.

内容审核大模型安全置信度校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。