发现主流毒性评测存在严重偏差,换任务类型就可能误判内容安全
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

- 测试时改变任务类型(如从续写变摘要),评测结果会显著变化
- 部分评测在不同数据领域下表现不一致,稳定性差
- 提醒开发者需构建更鲁棒的模型安全评估体系
大型语言模型在科研与产业中的快速应用凸显了其安全部署的挑战,暴露出毒性评测体系系统性评估的缺失。随着机构依赖这些评测来认证面向用户的模型及自动化内容审核系统,未被察觉的评测偏差可能导致脆弱或不安全系统的上线。本文探究现有评测框架的鲁棒性,分析被忽视的内在偏差,如模型选择、评估指标和任务类型的影响。实验显示,当评测任务从文本续写改为摘要生成时,评测系统对内容的有害性判断倾向显著增强;部分评测在输入数据领域变更后行为不一致。此外,还观察到模型特异性不稳定现象,表明亟需更稳健、全面的安全评估框架。
原文摘要 · Abstract (English)
The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluation of toxicity benchmarks. As organizations increasingly rely on these benchmarks to certify models for customer-facing applications and automated moderation, unrecognized evaluation biases could lead to the deployment of vulnerable or unsafe systems. This work investigates the robustness of established benchmarking setups and examines how to measure currently neglected intrinsic biases, such as those related to model choice, metrics, and task types. Our experiments uncover significant discrepancies in benchmark behaviors when evaluation setups are altered. Specifically, shifting the task from text completion to summarization increases the tendency of benchmarks to flag content as harmful. Additionally, certain benchmarks fail to maintain consistent behavior when the input data domain is changed. Furthermore, we observe model-specific instabilities, demonstrating a clear need for more robust and comprehensive safety evaluation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。