arXiv:2505.14534cs.CRcs.LG2025-05被引 40

测试并提升Gemini模型对抗恶意指令注入的防御能力。

Lessons from Defending Gemini Against Indirect Prompt Injections

  • 构建持续自适应攻击框架,模拟真实威胁环境。
  • 发现模型在未受保护时易受恶意数据操控,导致权限误用。
  • 适合关注大模型安全与可信部署的研究者与开发者。

Gemini 被越来越多用于代表用户执行任务,其函数调用和工具使用能力使其可访问用户数据。然而,部分工具需处理不受信任的数据,带来风险。攻击者可在这些数据中嵌入恶意指令,使模型偏离用户预期,误操作数据或权限。本文介绍 Google DeepMind 对 Gemini 模型对抗鲁棒性的评估方法,并总结关键经验。通过一个持续运行的对抗评估框架,部署一系列自适应攻击技术,对 Gemini 的过去、当前及未来版本进行持续测试。这些评估直接推动了模型在面对复杂攻击时的抗干扰能力提升。

原文摘要 · Abstract (English)

Gemini is increasingly used to perform tasks on behalf of users, where function-calling and tool-use capabilities enable the model to access user data. Some tools, however, require access to untrusted data introducing risk. Adversaries can embed malicious instructions in untrusted data which cause the model to deviate from the user's expectations and mishandle their data or permissions. In this report, we set out Google DeepMind's approach to evaluating the adversarial robustness of Gemini models and describe the main lessons learned from the process. We test how Gemini performs against a sophisticated adversary through an adversarial evaluation framework, which deploys a suite of adaptive attack techniques to run continuously against past, current, and future versions of Gemini. We describe how these ongoing evaluations directly help make Gemini more resilient against manipulation.

模型安全对抗攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。