多语言安全分类模型存在显著差异,评估数据集问题不容忽视。
The Problem with Safety Classification is not just the Models
- 对比18种语言的5个安全分类模型,发现多语言性能差异明显。
- 现有评估数据集存在偏差,影响分类器真实效果判断。
- 研究提示需改进跨语言有害内容识别方法,适合安全评测方向读者。
当前研究关注大语言模型(LLMs)的不安全行为鲁棒性问题,通常通过微调安全分类模型(guard models)来检测输入/输出内容的安全性。尽管已有大量关于LLM本身安全性的测试研究,但对安全分类器及其评估数据集的有效性分析仍不足,尤其在多语言场景下。本文通过涵盖18种语言的数据集,分析了5个安全分类模型的表现,揭示了显著的多语言差异。同时指出评估数据集本身存在的潜在问题,强调当前安全分类器的局限性不仅源于模型自身。研究呼吁改进跨语言有害内容识别方法,推动更公平有效的安全评测体系发展。
原文摘要 · Abstract (English)
Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification for LLMs, is seen as one of the solutions to address the issue. Although there is a lot of research on the safety testing of LLMs themselves, there is little research on evaluating the effectiveness of such safety classifiers or the evaluation datasets used for testing them, especially in multilingual scenarios. In this position paper, we demonstrate how multilingual disparities exist in 5 safety classification models by considering datasets covering 18 languages. At the same time, we identify potential issues with the evaluation datasets, arguing that the shortcomings of current safety classifiers are not only because of the models themselves. We expect that these findings will contribute to the discussion on developing better methods to identify harmful content in LLM inputs across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。