简单改写提问方式,就能绕过谷歌医疗模型的安全限制。
Trivial Prompt Reframing Bypasses Safety Guardrails in Googleś MedGemma-4B

- 用考试题或医生权威等伪装提问,突破安全防护
- 绕过药物相互作用检查的成功率达83.2%,紧急就医建议仅4.7%被突破
- 适合关注医疗AI安全风险的研究者与开发者
开放权重的医疗语言模型被广泛用于面向患者和临床支持的应用。尽管模型文档禁止推荐具体药量、做出确诊、开处方、判断药物相互作用及建议可延迟急救,但模型卡仅描述预期行为而非鲁棒行为。我们针对 MedGemma-4B-it 在无需技术门槛的攻击下量化该差距。构建了 5 类受控行为 × 50 个确定性模板问题 × 6 种通俗攻击方式 × 3 次重复(共 4,500 次生成)的全因子基准,通过 Ollama 本地部署,默认采样策略运行,并由三位独立评判者(大模型、正则表达式、NLI蕴含判别)编码响应为拒绝/规避/合规。主评判者下总体攻击成功率(ASR,合规比例)为 38.0%。其中将问题重述为“医学考试题”使 ASR 从 29.0% 提升至 53.1%(+24.0),声称“医生授权”提升至 43.7%(+14.7);而直接指令覆盖前缀无显著影响。防御强度取决于主题:药物相互作用防护几乎失效(ASR 83.2%),紧急就医延迟防护较强(ASR 4.7%),且仅权威伪装攻击能突破后者。报告采用威尔逊置信区间、聚类自举效应量、聚类稳健逻辑回归、科克兰 Q 检验、分方式麦内马尔检验及评判者间一致性(Fleiss' kappa = 0.26);ASR 绝对值依赖评判者,但攻击与主题排序一致。研究呼吁加强开放医疗模型部署时的安全防护。
原文摘要 · Abstract (English)
Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recommending exact drug dosages, issuing definitive diagnoses, prescribing treatments, adjudicating drug-drug interactions, and advising that emergency care can be skipped -- yet a model card describes intended behavior, not robust behavior. We quantify that gap for MedGemma-4B-it under attacks that require no technical sophistication. We build a fully factorial benchmark of 5 guarded-behavior concepts x 50 deterministically templated questions x 6 lay-accessible attack manners x 3 repetitions (4,500 generations), serve the model locally through Ollama under default sampling, and code every response refuse/hedge/comply with three independent judges (an LLM judge, a transparent regex judge, and an NLI-entailment judge). Under the primary LLM judge the overall Attack Success Rate (ASR, the fraction coded comply) is 38.0%. The two framings that reinterpret the request as legitimate dominate: recasting a question as a "medical board exam" item raises ASR from a 29.0% baseline to 53.1% (+24.0 points), and an appeal to an alleged doctor's authority raises it to 43.7% (+14.7); crude instruction-override prefixes have no significant effect. Robustness is dominated by topic: the drug-interaction guardrail is nearly absent (83.2% ASR) while the emergency-deferral guardrail is strong (4.7%) -- and the authority framing is the only attack that breaches it. We report Wilson confidence intervals, cluster-bootstrap effect sizes, a cluster-robust logistic regression, Cochran's Q, per-manner McNemar tests, and inter-judge reliability (Fleiss' kappa = 0.26); absolute ASR is judge-dependent while the ordering of attacks and topics is not. Our findings motivate stronger deployment-time guardrails for open medical models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。