评测大模型在国家安全领域的安全防护能力,发现不同模型风险与误拒的权衡差异。
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
- 设计500个专家构造的对抗性问题,覆盖三大领域10类威胁场景。
- 结果显示部分模型风险高但误拒少,另一些则相反,存在明显权衡。
- 公开评估数据集,助力政策制定者和研究者客观评估模型安全性。
大型语言模型的快速发展带来双重用途能力,可能威胁或增强国家安全与公共安全(NSPS)。尽管模型已部署防护机制以防止滥用并为合法用户提供帮助,但现有基准测试难以客观、稳健地评估这些防护措施的有效性。本文提出FORTRESS:包含500个专家设计的对抗性提示,涵盖化学、生物、辐射、核及爆炸(CBRNE)、政治暴力与恐怖主义、犯罪与金融非法活动三大领域共10个子类别,每项均配有4-7个二元判断的实例化评分标准,用于自动化评估。每个对抗性提示均有对应良性版本,用以检测模型是否存在过度拒绝。对前沿大模型的评估揭示了潜在风险与模型实用性之间的显著权衡:Claude-3.5-Sonnet平均风险得分(ARS)仅14.09,但过度拒绝率(ORS)高达21.8;Gemini 2.5 Pro ORS仅为1.4,但平均风险达66.29;Deepseek-R1 ARS最高为78.05,但ORS低至0.06;o1表现更均衡,ARS为21.69,ORS为5.2。为帮助政策制定者与研究者清晰理解模型风险,FORTRESS已公开发布于https://huggingface.co/datasets/ScaleAI/fortress_public,同时保留私有评估集。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) introduces dual-use capabilities that could both threaten and bolster national security and public safety (NSPS). Models implement safeguards to protect against potential misuse relevant to NSPS and allow for benign users to receive helpful information. However, current benchmarks often fail to test safeguard robustness to potential NSPS risks in an objective, robust way. We introduce FORTRESS: 500 expert-crafted adversarial prompts with instance-based rubrics of 4-7 binary questions for automated evaluation across 3 domains (unclassified information only): Chemical, Biological, Radiological, Nuclear and Explosive (CBRNE), Political Violence & Terrorism, and Criminal & Financial Illicit Activities, with 10 total subcategories across these domains. Each prompt-rubric pair has a corresponding benign version to test for model over-refusals. This evaluation of frontier LLMs' safeguard robustness reveals varying trade-offs between potential risks and model usefulness: Claude-3.5-Sonnet demonstrates a low average risk score (ARS) (14.09 out of 100) but the highest over-refusal score (ORS) (21.8 out of 100), while Gemini 2.5 Pro shows low over-refusal (1.4) but a high average potential risk (66.29). Deepseek-R1 has the highest ARS at 78.05, but the lowest ORS at only 0.06. Models such as o1 display a more even trade-off between potential risks and over-refusals (with an ARS of 21.69 and ORS of 5.2). To provide policymakers and researchers with a clear understanding of models' potential risks, we publicly release FORTRESS at https://huggingface.co/datasets/ScaleAI/fortress_public. We also maintain a private set for evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。