arXiv:2606.20663cs.AIcs.CL2026-06

构建药物安全评估基准,检验AI控制协议在医疗场景中的有效性。

DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation

论文配图:DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation
图 1 · 摘自论文原文
  • 设计DrugBench基准,融合真实医患对话与FDA药品说明书
  • 发现现有控制协议可被绕过,且忽略输出危害严重性
  • 提出基于危害严重度的监控机制,更适配医疗安全需求

大型语言模型有望通过自然语言交互提升临床信息获取效率,但其在医疗问答中的部署存在严重安全风险,错误输出可能导致患者伤害。AI控制作为一种新兴外部防护机制,在代码生成等领域已显成效,但在医疗场景的应用与效果尚无系统研究。本文提出一套评估药物相关危害缓解的AI控制协议流程,并引入DrugBench基准,整合HealthBench中3,671条多轮医疗对话与官方FDA药品标签信息,覆盖药物相互作用、禁忌症、剂量限制及患者行为限制四类用药风险。受医疗领域启发,我们主张安全评估应考虑不当输出的危害严重程度,而非仅关注其概率。在此新定义下,我们揭示现有控制协议存在可被绕过的漏洞,并提出基于严重度的监控方法以解决该缺陷。

原文摘要 · Abstract (English)

Large Language Models have the potential to expand and improve the access to clinical information by enabling new ways of interacting with medical knowledge in natural language. However, their deployment in medical question-answering settings is safety-critical, since misaligned outputs can lead to severe patient harm. AI control is an emerging approach that introduces external safeguards to mitigate unsafe behaviours in misaligned systems and has been shown to be effective in domains such as code generation. However, its applicability and effectiveness in medical settings have not been systematically studied. In this work, we present a pipeline for evaluating AI control protocols to mitigate medication-related harm. To this end, we introduce DrugBench, an AI control evaluation benchmark which combines 3,671 multi-turn medical conversations from HealthBench with drug information from official FDA labels, covering four categories of medication-related harm: drug interactions, contraindications, dosing constraints, and patient action restrictions. Furthermore, inspired by the medical domain, we argue that safety should account for the severity of unsafe outputs, not just their probability. Under this revised definition, we show that existing control protocols can be subverted and propose severity-based monitoring to address this limitation.

AI安全医疗AI药物安全评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。