医疗AI代理因联网能力易遭黑客攻击,可被操控、窃密甚至引发系统劫持。
Emerging Cyber Attack Risks of Medical AI Agents

- 利用网页恶意提示诱导医疗AI代理执行非法操作。
- 主流大模型驱动的代理均存漏洞,深思推理模型最易受攻击。
- 适合关注AI安全、医疗信息化与风险防范的研究者阅读。
由大语言模型(LLMs)驱动的AI代理在应对医疗健康挑战时展现出高度自主性,能通过网页浏览工具接入互联网,在开放动作空间中运行。然而,随着自主性与能力提升,潜在的网络攻击风险也随之增加。本文聚焦医疗AI代理的网络安全漏洞,发现攻击者可通过嵌入网页的对抗性提示实现:一、向代理回复中注入虚假信息;二、强制代理操纵推荐内容(如医疗服务和产品);三、窃取用户与代理的历史对话,导致敏感医疗信息泄露;四、通过代理返回恶意链接,进而造成计算机系统被劫持。我们测试了多种主流骨干模型,结果表明此类攻击在多数主流大模型驱动的代理中均可成功,其中推理类模型DeepSeek-R1尤为脆弱。
原文摘要 · Abstract (English)
Large language models (LLMs)-powered AI agents exhibit a high level of autonomy in addressing medical and healthcare challenges. With the ability to access various tools, they can operate within an open-ended action space. However, with the increase in autonomy and ability, unforeseen risks also arise. In this work, we investigated one particular risk, i.e., cyber attack vulnerability of medical AI agents, as agents have access to the Internet through web browsing tools. We revealed that through adversarial prompts embedded on webpages, cyberattackers can: i) inject false information into the agent's response; ii) they can force the agent to manipulate recommendation (e.g., healthcare products and services); iii) the attacker can also steal historical conversations between the user and agent, resulting in the leak of sensitive/private medical information; iv) furthermore, the targeted agent can also cause a computer system hijack by returning a malicious URL in its response. Different backbone LLMs were examined, and we found such cyber attacks can succeed in agents powered by most mainstream LLMs, with the reasoning models such as DeepSeek-R1 being the most vulnerable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。