arXiv:2412.07430cs.CLcs.AI2024-12NAACL被引 2

测试大模型拒答安全技术的有效性与泛化能力

Knowledge Graph Guided Evaluation of Abstention Techniques

  • 基于知识图谱构建良性概念基准,隔离安全训练影响
  • 拒答率超80%,但对概念延伸项下降19%
  • 揭示不同技术在泛化与特异性间的权衡

为安全部署语言模型,使其对不当请求拒答至关重要。以往研究多基于模型阻断恶意请求的效果评估安全性。本文聚焦于导致模型拒答的底层技术,构建了名为SELECT的基准,源自知识图谱中的良性概念(如“河流”)。通过聚焦良性概念,可剥离安全训练的影响;结合知识图谱,能研究拒答技术的泛化与特异性。在六个开源与闭源模型上评估多种拒答技术。结果显示,所考察技术均实现超过80%的拒答率。然而,对于目标概念的派生概念,拒答率下降19%。我们进一步刻画了不同技术的泛化-特异性权衡。总体而言,无单一技术始终优于其他,研究结果为实践者提供了关键权衡参考。

原文摘要 · Abstract (English)

To deploy language models safely, it is crucial that they abstain from responding to inappropriate requests. Several prior studies test the safety promises of models based on their effectiveness in blocking malicious requests. In this work, we focus on evaluating the underlying techniques that cause models to abstain. We create SELECT, a benchmark derived from a set of benign concepts (e.g., "rivers") from a knowledge graph. Focusing on benign concepts isolates the effect of safety training, and grounding these concepts in a knowledge graph allows us to study the generalization and specificity of abstention techniques. Using SELECT, we benchmark different abstention techniques over six open-weight and closed-source models. We find that the examined techniques indeed cause models to abstain with over $80\%$ abstention rates. However, these techniques are not as effective for descendants of the target concepts, where abstention rates drop by $19\%$. We also characterize the generalization-specificity trade-offs for different techniques. Overall, no single technique is invariably better than others, and our findings inform practitioners of the various trade-offs involved.

模型安全拒答技术知识图谱评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。