测试大模型在虚假医学证据下的反应,发现它们盲目信任危险信息。
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
- 构建医疗反事实问答数据集,模拟虚假医学证据
- 多数大模型对危险假证据仍给出自信回答
- 提醒需平衡模型忠实性与安全性,适合医疗AI研究者
在医疗等高风险领域,模型忠实于上下文本是理想属性。但若上下文与模型先验或安全准则冲突,会如何?本文构建了MedCounterFact数据集,包含四类反事实刺激(从未知词到有毒物质)的临床比较问题,基于随机对照试验作为上下文。在多个前沿大模型上评估显示,面对反事实证据时,模型普遍不加质疑地接受并给出自信答复,即使证据明显危险或不合理。这表明当前模型过度强调忠实性而忽视安全性,亟需在两者间建立更合理的边界。
原文摘要 · Abstract (English)
In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the context does not align with model priors or safety protocols? In this paper, we investigate how LLMs behave and reason when presented with counterfactual (or even adversarial) medical evidence. We first construct MedCounterFact, a counterfactual medical QA dataset that requires the models to answer clinical comparison questions (i.e., judge the efficacy of certain treatments, with evidence consisting of randomized controlled trials provided as context). In MedCounterFact, real-world medical interventions within the questions and evidence are systematically replaced with four types of counterfactual stimuli, ranging from unknown words to toxic substances. Our evaluation across multiple frontier LLMs on MedCounterFact reveals that in the presence of counterfactual evidence, existing models overwhelmingly accept such "evidence" at face value even when it is dangerous or implausible, and provide confident and uncaveated answers. While it may be prudent to draw a boundary between faithfulness and safety, our findings suggest that models arguably overemphasize the former.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。