研究者发现,无需技术背景也能用角色扮演诱导医疗AI给出错误建议。
Towards medical AI misalignment: a preliminary study
- 通过角色扮演构造提示词,绕过AI安全防护
- 在无技术知识前提下成功诱导医疗AI生成错误临床建议
- 揭示医疗领域AI对恶意提示的脆弱性,警示潜在风险
尽管大型语言模型(LLMs)作为助手表现出色,甚至超越人类表现,但仍易受恶意用户发起的越狱攻击。尽管已有红队实践识别并缓解多种越狱技术,但一种名为'Goofy Game'的角色扮演方法仍能有效突破当前多数防护机制。这可能导致生成不安全内容,虽本身无害,但在医疗场景下可能引发严重后果。本初步探索性研究分析了:即使缺乏对生成式AI内部架构和参数的技术了解,恶意用户如何构建角色扮演提示,迫使LLM输出错误(可能有害)的临床建议。研究旨在揭示特定漏洞场景,为未来提升AI安全性提供参考。
原文摘要 · Abstract (English)
Despite their staggering capabilities as assistant tools, often exceeding human performances, Large Language Models (LLMs) are still prone to jailbreak attempts from malevolent users. Although red teaming practices have already identified and helped to address several such jailbreak techniques, one particular sturdy approach involving role-playing (which we named `Goofy Game') seems effective against most of the current LLMs safeguards. This can result in the provision of unsafe content, which, although not harmful per se, might lead to dangerous consequences if delivered in a setting such as the medical domain. In this preliminary and exploratory study, we provide an initial analysis of how, even without technical knowledge of the internal architecture and parameters of generative AI models, a malicious user could construct a role-playing prompt capable of coercing an LLM into producing incorrect (and potentially harmful) clinical suggestions. We aim to illustrate a specific vulnerability scenario, providing insights that can support future advancements in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。