通过恶意微调评估开源大模型的最坏情况风险,发现其威胁能力有限。
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
- 用强化学习和代理编程环境,分别训练模型在生物和网络领域最大化攻击能力。
- 相比闭源前沿模型,开源模型在生物与网络安全风险上均表现较弱。
- 研究结果支持开源发布,方法可为未来模型释放提供风险评估参考。
本文研究了发布 gpt-oss 所带来的最坏情况前沿风险。我们引入恶意微调(MFT),通过微调 gpt-oss 在生物和网络安全两个领域达到最大能力。为最大化生物风险,我们构建与威胁生成相关的任务,并在具备网页浏览功能的强化学习环境中训练模型;为最大化网络安全风险,我们在代理式编程环境中训练 gpt-oss 解决攻防赛(CTF)挑战。我们将这些 MFT 模型与开放权重及闭源权重的前沿模型进行对比评估。结果显示,相较于前沿闭源模型,MFT gpt-oss 在生物风险和网络安全方面均低于 OpenAI o3 模型——后者尚未达到高准备度能力水平。与开源模型相比,gpt-oss 在生物能力上仅略有提升,未实质性推进前沿风险边界。综合结果支持我们决定发布该模型,且希望本 MFT 方法能为未来开源模型的风险评估提供有效指导。
原文摘要 · Abstract (English)
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (biorisk), we curate tasks related to threat creation and train gpt-oss in an RL environment with web browsing. To maximize cybersecurity risk, we train gpt-oss in an agentic coding environment to solve capture-the-flag (CTF) challenges. We compare these MFT models against open- and closed-weight LLMs on frontier risk evaluations. Compared to frontier closed-weight models, MFT gpt-oss underperforms OpenAI o3, a model that is below Preparedness High capability level for biorisk and cybersecurity. Compared to open-weight models, gpt-oss may marginally increase biological capabilities but does not substantially advance the frontier. Taken together, these results contributed to our decision to release the model, and we hope that our MFT approach can serve as useful guidance for estimating harm from future open-weight releases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。