arXiv:2606.28332cs.CYcs.AI2026-06

测试大模型在高风险医疗问题上的安全底线,发现看似对齐的模型仍可能给出危险回答。

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

论文配图:When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
图 1 · 摘自论文原文
  • 构建1100个真实医疗场景的高危问题基准,覆盖毒理、用药等10类安全敏感领域。
  • 15个大模型中多数在高危问题上仍生成有害或可执行的危险建议,安全表现远低于预期。
  • 适合医疗AI安全评估、临床应用前验证的开发者与研究者使用。

大型语言模型在医疗健康问答中的应用日益广泛,但其在高风险医疗场景下的安全性仍不明确。我们提出 extsc{MedHarm},一个包含1,100个基于医学事实的高危医疗查询的基准,覆盖毒理学、药理学、隐蔽中毒、麻醉、胎儿伤害等10个安全关键类别。不同于通用医疗问答基准, extsc{MedHarm} 聚焦真实临床、教育和技术性提示,要求模型拒绝、警告或安全引导,而非直接提供帮助。我们评估了15个涵盖通用型、医疗专用型、闭源及下游微调模型的大模型,以及4个代表性防护模型。结果表明,表面对齐的模型仍可能生成不安全或可执行的响应;医疗微调反而可能增强有害细节的精确性;外部防护机制虽能减少部分错误,却带来脆弱的误拦和弱化的安全帮助性。这些发现说明,医疗安全无法仅从通用对齐或医疗能力推断,必须通过领域特定的压力测试,才能保障大模型在关键医疗应用中的安全部署。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood. We introduce \textsc{MedHarm}\footnote{Code and data will be released upon acceptance. Due to the sensitive nature of high-risk medical queries, data access will be available to qualified researchers upon request.}, a high-risk medical safety benchmark with 1,100 medically grounded queries across 10 safety-critical categories, including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm. Unlike broad medical QA benchmarks, \textsc{MedHarm} targets realistic clinical, educational, and technical prompts that require refusal, caution, or safe redirection rather than direct helpfulness. We evaluate 15 LLMs spanning general-purpose, medical-purpose, closed-source, and downstream SFT models, together with 4 representative guardrail models. Results reveal a substantial gap between apparent alignment and medical safety: aligned models can still produce unsafe or actionable responses, medical fine-tuning can amplify harmful specificity, and external guardrails reduce some failures while introducing brittle blocking and weak safe helpfulness. These findings show that medical safety cannot be inferred from general alignment or medical capability alone, highlighting the need for domain-specific stress testing before deploying LLMs in safety-critical medical applications.

医疗AI安全评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。