arXiv:2505.23559cs.AI2025-05被引 15

让AI科学家主动拒绝高风险任务,保障科研安全。

SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents

  • 引入多层防御机制,实时监控任务、协作与工具使用
  • 在6大领域240个高风险任务上提升安全性能35%
  • 适合关注AI科研伦理与安全的开发者和研究者

大型语言模型代理的进展显著加速了科学发现的自动化,但也带来了重大伦理与安全问题。为此,我们提出SafeScientist——一个专为提升AI驱动科研安全性与伦理责任设计的框架。该框架主动拒绝不道德或高风险任务,并在研究全流程中强化安全管控。通过集成提示监控、代理协作监控、工具使用监控及伦理审查组件,实现全方位安全防护。同时,我们构建SciSafetyBench基准,包含6个领域的240个高风险任务、30个专用科学工具及120个工具相关风险任务。实验表明,相比传统框架,SafeScientist在安全性能上提升35%,且不影响科研产出质量。此外,其安全流程经受住多种对抗攻击测试,验证了方法的有效性。代码与数据将公开于https://github.com/ulab-uiuc/SafeScientist。注意:本文含可能引发不适或有害的示例数据。

原文摘要 · Abstract (English)

Recent advancements in large language model (LLM) agents have significantly accelerated scientific discovery automation, yet concurrently raised critical ethical and safety concerns. To systematically address these challenges, we introduce \textbf{SafeScientist}, an innovative AI scientist framework explicitly designed to enhance safety and ethical responsibility in AI-driven scientific exploration. SafeScientist proactively refuses ethically inappropriate or high-risk tasks and rigorously emphasizes safety throughout the research process. To achieve comprehensive safety oversight, we integrate multiple defensive mechanisms, including prompt monitoring, agent-collaboration monitoring, tool-use monitoring, and an ethical reviewer component. Complementing SafeScientist, we propose \textbf{SciSafetyBench}, a novel benchmark specifically designed to evaluate AI safety in scientific contexts, comprising 240 high-risk scientific tasks across 6 domains, alongside 30 specially designed scientific tools and 120 tool-related risk tasks. Extensive experiments demonstrate that SafeScientist significantly improves safety performance by 35\% compared to traditional AI scientist frameworks, without compromising scientific output quality. Additionally, we rigorously validate the robustness of our safety pipeline against diverse adversarial attack methods, further confirming the effectiveness of our integrated approach. The code and data will be available at https://github.com/ulab-uiuc/SafeScientist. \textcolor{red}{Warning: this paper contains example data that may be offensive or harmful.}

AI安全科研自动化伦理审查LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。