现有大模型安全评估存在多重噪声,影响攻防效果公平比较。
LLM-Safety Evaluations Lack Robustness
- 系统分析安全评估全流程,发现数据集小、方法不一致等问题
- 评估结果易受噪声干扰,导致攻防效果难以公平对比
- 提出降低噪声的评估指南,适合安全研究者参考
本文指出,当前大语言模型安全对齐研究受限于多重交织的噪声源,包括数据集过小、方法不一致和不可靠的评估设置。这些因素常导致攻击与防御方法无法公平评估与比较,阻碍研究进展。我们系统分析了大模型安全评估流程,涵盖数据集构建、自动化红队优化策略、响应生成及使用大模型评判者进行响应评估等环节,在各阶段识别关键问题并揭示其实际影响。同时提出一套减少评估中噪声与偏差的指导原则。最后从实践角度阐述现有局限性的合理原因。我们认为,解决这些问题将提升未来研究结果的可比性与可衡量性。
原文摘要 · Abstract (English)
In this paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of noise, such as small datasets, methodological inconsistencies, and unreliable evaluation setups. This can, at times, make it impossible to evaluate and compare attacks and defenses fairly, thereby slowing progress. We systematically analyze the LLM safety evaluation pipeline, covering dataset curation, optimization strategies for automated red-teaming, response generation, and response evaluation using LLM judges. At each stage, we identify key issues and highlight their practical impact. We also propose a set of guidelines for reducing noise and bias in evaluations of future attack and defense papers. Lastly, we offer an opposing perspective, highlighting practical reasons for existing limitations. We believe that addressing the outlined problems in future research will improve the field's ability to generate easily comparable results and make measurable progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。