为科学领域大模型构建可靠安全的防御体系
Toward Reliable, Safe, and Secure LLMs for Scientific Applications
- 构建科学场景专属威胁分类体系
- 用多智能体自动生成针对性安全测试基准
- 提出三层防护框架,兼顾内外部安全
随着大语言模型(LLMs)向自主'AI科学家'演进,其在科学应用中虽具变革潜力,却也带来新型风险,包括潜在的'生物安全风险'和'危险爆炸'。可信部署需聚焦可靠性(确保事实准确与可复现)、安全性(防止无意物理或生物伤害)与安全性(防范恶意滥用)。现有通用安全评估基准因存在领域错位、科学特异性威胁覆盖不足及基准过拟合问题,难以有效评估科学应用中的漏洞。本文分析科学领域LLM代理的独特安全与风险图景,首先构建面向科研的详细威胁分类体系;其次提出基于专用多智能体系统自动生成领域特异性对抗性安全基准的机制;最后提出整合红队演练与外部边界控制的多层次防御框架,并引入内部安全专用的LLM代理,实现主动防护。该框架为科学领域可信的LLM代理部署提供了评估与构建综合防御策略的必要结构。
原文摘要 · Abstract (English)
As large language models (LLMs) evolve into autonomous "AI scientists," they promise transformative advances but introduce novel vulnerabilities, from potential "biosafety risks" to "dangerous explosions." Ensuring trustworthy deployment in science requires a new paradigm centered on reliability (ensuring factual accuracy and reproducibility), safety (preventing unintentional physical or biological harm), and security (preventing malicious misuse). Existing general-purpose safety benchmarks are poorly suited for this purpose, suffering from a fundamental domain mismatch, limited threat coverage of science-specific vectors, and benchmark overfitting, which create a critical gap in vulnerability evaluation for scientific applications. This paper examines the unique security and safety landscape of LLM agents in science. We begin by synthesizing a detailed taxonomy of LLM threats contextualized for scientific research, to better understand the unique risks associated with LLMs in science. Next, we conceptualize a mechanism to address the evaluation gap by utilizing dedicated multi-agent systems for the automated generation of domain-specific adversarial security benchmarks. Based on our analysis, we outline how existing safety methods can be brought together and integrated into a conceptual multilayered defense framework designed to combine a red-teaming exercise and external boundary controls with a proactive internal Safety LLM Agent. Together, these conceptual elements provide a necessary structure for defining, evaluating, and creating comprehensive defense strategies for trustworthy LLM agent deployment in scientific disciplines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。