构建医疗大模型安全测试数据集,识别真实场景下的伦理风险
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
- 基于巴西医保系统设计21万条对抗性提示,覆盖7类医疗违规行为
- 每条提示附带违规理由,帮助模型理解伦理边界而非机械拒绝
- 适合医疗AI安全研究者、监管机构及高风险场景系统开发者
将大语言模型(LLMs)融入医疗领域需以‘首要不伤害’为安全准则。然而现有对齐技术依赖通用伤害定义,无法捕捉如行政欺诈和临床歧视等情境相关违规行为。为此,我们提出Medical Malice:一个包含214,219条对抗性提示的数据集,其校准于巴西统一健康系统(SUS)的法规与伦理复杂性。关键在于,每条提示均附有违规原因,使模型能内化伦理边界,而非仅记忆固定拒答。通过在人格驱动流程中使用未对齐代理(Grok-4),我们合成七类高保真威胁,涵盖采购操纵、排队插队至产科暴力。我们讨论释放这些“漏洞签名”的伦理设计,以弥合恶意使用者与AI开发者间的信息不对称。最终,本文倡导从普适性安全转向情境感知安全,提供免疫医疗AI应对高风险医疗环境中细微且系统性威胁所需资源——此类漏洞是患者安全与医疗AI成功集成的最大风险。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into healthcare demands a safety paradigm rooted in \textit{primum non nocere}. However, current alignment techniques rely on generic definitions of harm that fail to capture context-dependent violations, such as administrative fraud and clinical discrimination. To address this, we introduce Medical Malice: a dataset of 214,219 adversarial prompts calibrated to the regulatory and ethical complexities of the Brazilian Unified Health System (SUS). Crucially, the dataset includes the reasoning behind each violation, enabling models to internalize ethical boundaries rather than merely memorizing a fixed set of refusals. Using an unaligned agent (Grok-4) within a persona-driven pipeline, we synthesized high-fidelity threats across seven taxonomies, ranging from procurement manipulation and queue-jumping to obstetric violence. We discuss the ethical design of releasing these "vulnerability signatures" to correct the information asymmetry between malicious actors and AI developers. Ultimately, this work advocates for a shift from universal to context-aware safety, providing the necessary resources to immunize healthcare AI against the nuanced, systemic threats inherent to high-stakes medical environments -- vulnerabilities that represent the paramount risk to patient safety and the successful integration of AI in healthcare systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。