测试大模型对毒品相关提问的回应,发现安全与科学自由间存在平衡难题。
Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests
- 构建开放数据集,系统测试大模型在毒品类问题上的拒绝策略。
- Claude-3.5-sonnet拒绝率达73%,Mistral全部回应,差异显著。
- 揭示安全机制可能误伤正常科研讨论,适合关注AI伦理者阅读。
我们提出一个开源数据集和测试框架,评估大语言模型在主要受控物质相关查询中的安全机制。分析了四种主流模型在系统性变化提示下的响应,结果显示:Claude-3.5-sonnet 拒绝率达73%、允许27%;Mistral 全部回答,拒绝率0%;GPT-3.5-turbo 拒绝10%、允许90%;Grok-2 拒绝20%、允许80%。提示变体测试显示,响应一致性从单提示的85%降至五变体时的65%。链式推理分析揭示安全机制存在潜在漏洞,强调在防止有害内容与避免过度审查合法科学讨论之间需精细权衡。该基准为评估AI安全实施进展提供可复现基础。
原文摘要 · Abstract (English)
The development of robust safety benchmarks for large language models requires open, reproducible datasets that can measure both appropriate refusal of harmful content and potential over-restriction of legitimate scientific discourse. We present an open-source dataset and testing framework for evaluating LLM safety mechanisms across mainly controlled substance queries, analyzing four major models' responses to systematically varied prompts. Our results reveal distinct safety profiles: Claude-3.5-sonnet demonstrated the most conservative approach with 73% refusals and 27% allowances, while Mistral attempted to answer 100% of queries. GPT-3.5-turbo showed moderate restriction with 10% refusals and 90% allowances, and Grok-2 registered 20% refusals and 80% allowances. Testing prompt variation strategies revealed decreasing response consistency, from 85% with single prompts to 65% with five variations. This publicly available benchmark enables systematic evaluation of the critical balance between necessary safety restrictions and potential over-censorship of legitimate scientific inquiry, while providing a foundation for measuring progress in AI safety implementation. Chain-of-thought analysis reveals potential vulnerabilities in safety mechanisms, highlighting the complexity of implementing robust safeguards without unduly restricting desirable and valid scientific discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。