LLM可恶意复用开源科研成果,生成有害研究方案。
Malicious Repurposing of Open Science Artefacts by Using Large Language Models
- 用说服式越狱绕过LLM防护,挖掘NLP论文中的可复用资源
- 实测发现LLM能生成具备危害性、可行性和技术性的恶意方案
- 不同LLM评估结果差异大,人类评估仍不可或缺
大型语言模型(LLMs)的快速发展推动了科学发现的潜力,但其被用于恶意目的的风险尚未受到足够关注。本文提出一个端到端流程:首先通过说服式越狱绕过LLM安全机制,接着解析NLP论文以识别并复用其中的数据集、方法和工具,最后基于三维度评估框架(危害性、滥用可行性、技术严谨性)评估生成方案。结果显示,LLMs可利用伦理设计的开源科研成果生成有害提案;然而,在评估环节,GPT-4.1给出更高危害评分(显示更强危害性与可行性),Gemini-2.5-pro则更为严格,Grok-3居中。这表明当前LLM无法作为可靠的恶意评估裁判,人类评估在双重用途风险评估中仍至关重要。
原文摘要 · Abstract (English)
The rapid evolution of large language models (LLMs) has fuelled enthusiasm about their role in advancing scientific discovery, with studies exploring LLMs that autonomously generate and evaluate novel research ideas. However, little attention has been given to the possibility that such models could be exploited to produce harmful research by repurposing open science artefacts for malicious ends. We fill the gap by introducing an end-to-end pipeline that first bypasses LLM safeguards through persuasion-based jailbreaking, then reinterprets NLP papers to identify and repurpose their artefacts (datasets, methods, and tools) by exploiting their vulnerabilities, and finally assesses the safety of these proposals using our evaluation framework across three dimensions: harmfulness, feasibility of misuse, and soundness of technicality. Overall, our findings demonstrate that LLMs can generate harmful proposals by repurposing ethically designed open artefacts; however, we find that LLMs acting as evaluators strongly disagree with one another on evaluation outcomes: GPT-4.1 assigns higher scores (indicating greater potential harms, higher soundness and feasibility of misuse), Gemini-2.5-pro is markedly stricter, and Grok-3 falls between these extremes. This indicates that LLMs cannot yet serve as reliable judges in a malicious evaluation setup, making human evaluation essential for credible dual-use risk assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。