arXiv:2506.11094cs.CLcs.AI2025-06综述被引 51

系统梳理大模型安全评估的四大维度,为可靠部署提供指南

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

  • 构建四维评估框架:为何评、评什么、在哪评、如何评
  • 涵盖毒性、偏见、真实性等六大核心安全维度
  • 适合研究者与从业者快速掌握安全评估全貌

随着人工智能快速发展,大语言模型(LLMs)在自然语言处理领域展现出强大能力,包括内容生成、人机交互、机器翻译和代码生成等。然而其广泛应用也引发严重安全问题,生成内容可能表现出毒性、偏见或虚假信息,尤其在对抗性场景下,引起学术界与产业界的广泛关注。尽管已有诸多研究尝试评估这些风险,但针对大模型安全评估的系统性综述仍较缺乏。本文旨在填补这一空白,提出一个结构化的近期进展概述。具体地,我们构建了四维分类体系:(i) 为何评估,探讨安全评估的背景、与通用评估的区别及其重要性;(ii) 评什么,基于关键能力对现有安全评估任务进行分类,涵盖毒性、鲁棒性、伦理、偏见与公平性、真实性等维度;(iii) 在哪评,总结当前使用的评估指标、数据集与基准;(iv) 如何评,回顾主流评估方法,依据评估者角色及整合评估流程的框架。最后,我们识别当前挑战并提出未来研究方向,强调优先推进安全评估以确保大模型在真实应用中的可靠与负责任部署。

原文摘要 · Abstract (English)

With the rapid advancement of artificial intelligence, Large Language Models (LLMs) have shown remarkable capabilities in Natural Language Processing (NLP), including content generation, human-computer interaction, machine translation, and code generation. However, their widespread deployment has also raised significant safety concerns. In particular, LLM-generated content can exhibit unsafe behaviors such as toxicity, bias, or misinformation, especially in adversarial contexts, which has attracted increasing attention from both academia and industry. Although numerous studies have attempted to evaluate these risks, a comprehensive and systematic survey on safety evaluation of LLMs is still lacking. This work aims to fill this gap by presenting a structured overview of recent advances in safety evaluation of LLMs. Specifically, we propose a four-dimensional taxonomy: (i) Why to evaluate, which explores the background of safety evaluation of LLMs, how they differ from general LLMs evaluation, and the significance of such evaluation; (ii) What to evaluate, which examines and categorizes existing safety evaluation tasks based on key capabilities, including dimensions such as toxicity, robustness, ethics, bias and fairness, truthfulness, and related aspects; (iii) Where to evaluate, which summarizes the evaluation metrics, datasets and benchmarks currently used in safety evaluations; (iv) How to evaluate, which reviews existing mainstream evaluation methods based on the roles of the evaluators and some evaluation frameworks that integrate the entire evaluation pipeline. Finally, we identify the challenges in safety evaluation of LLMs and propose promising research directions to promote further advancement in this field. We emphasize the necessity of prioritizing safety evaluation to ensure the reliable and responsible deployment of LLMs in real-world applications.

大模型安全评估框架综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。