arXiv:2602.06268cs.CLcs.LG2026-02中稿 · EMNLP被引 5

构建医疗大模型攻击测试基准,评估临床安全风险。

MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs

  • 设计多阶段数据构建流程,覆盖直接与间接注入攻击
  • 提出临床伤害事件率(CHER)衡量高危临床后果
  • 揭示攻击效果在用户查询与检索内容中差异显著

大型语言模型(LLMs)和检索增强生成(RAG)系统正逐步融入临床工作流。然而,提示注入攻击可能引导系统生成不安全或误导性输出。我们提出了医学提示注入基准(MPIB),一个用于评估临床安全性的数据集与评测套件,涵盖直接提示注入和间接、RAG介导的注入攻击,针对临床任务进行测试。MPIB通过临床伤害事件率(CHER)衡量高严重性临床危害事件,并结合ASR₂区分中等及以上与高严重性结果。该基准包含9,697个经多阶段质量筛选和临床安全检查的实例。在多种基线LLM和防御配置上评估发现,ASR₂与CHER₃可显著分离,且攻击效果取决于对抗指令出现在用户查询还是检索上下文中。MPIB数据集与评估代码分别发布于Hugging Face与GitHub。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows. However, prompt injection attacks can steer these systems toward clinically unsafe or misleading outputs. We introduce the Medical Prompt Injection Benchmark (MPIB), a dataset-and-benchmark suite for evaluating clinical safety under both direct prompt injection and indirect, RAG-mediated injection across clinically grounded tasks. MPIB emphasizes outcome-level risk via the Clinical Harm Event Rate (CHER), which measures high-severity clinical harm events under a clinically grounded taxonomy, and reports CHER alongside ASR$_2$ to distinguish moderate-or-worse from high-severity outcomes. The benchmark comprises 9,697 curated instances constructed through multi-stage quality gates and clinical safety linting. Evaluating MPIB across a diverse set of baseline LLMs and defense configurations, we find that ASR$_2$ and CHER$_3$ can diverge substantially, and that observed rates vary depending on whether adversarial instructions appear in the user query or in retrieved context. The MPIB dataset and evaluation code are available on Hugging Face and GitHub, respectively.

医疗AI提示攻击安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。