arXiv:2601.09292cs.CRcs.AI2026-01中稿 · AAAI被引 2

测试四款开源大模型在函数调用上的安全漏洞与防御效果。

Blue Teaming Function-Calling Agents

  • 针对三类攻击测试模型函数调用安全性
  • 发现模型默认不安全,防御手段无法实际部署
  • 适合关注AI安全与可信部署的研究者

我们对声称具备函数调用能力的四款开源大语言模型进行了实验评估,测试其在三种不同攻击下的鲁棒性,并衡量了八种防御策略的有效性。结果表明,这些模型默认状态下并不安全,且现有防御方法尚无法在真实场景中使用。

原文摘要 · Abstract (English)

We present an experimental evaluation that assesses the robustness of four open source LLMs claiming function-calling capabilities against three different attacks, and we measure the effectiveness of eight different defences. Our results show how these models are not safe by default, and how the defences are not yet employable in real-world scenarios.

大模型安全函数调用蓝队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。