评估医疗大模型幻觉时,关注内容潜在危害性比只看真假更有意义。
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
- 用风险语言(如用药指令、紧急提示)量化幻觉危害程度。
- 不同模型表面表现相似,但实际风险差异显著。
- 适合关注医疗AI安全性的研究人员和临床应用开发者。
大型语言模型在面向患者的医疗问答中应用日益广泛,其生成的幻觉内容可能带来不同程度的伤害。然而现有幻觉评估标准和指标主要关注事实正确性,将所有错误视为同等严重,忽略了临床相关的失效模式,尤其是模型生成无依据却具有行动指导性的医疗建议。本文提出一种风险敏感的评估框架,通过识别治疗指令、禁忌症、紧急信号及高风险药物提及等风险承载语言来量化幻觉的危害性。该方法不评估临床正确性,而是评估若患者采取这些幻觉内容可能带来的潜在影响。进一步结合相关性评分,可识别出高风险且低依据支撑的失败案例。我们在三个指令微调模型上使用设计为安全压力测试的患者导向提示进行了测试,结果表明:尽管模型在表面行为上相似,其风险特征却存在显著差异;而传统评估指标无法捕捉这些关键区别。研究强调了在幻觉评估中引入风险敏感性的重要性,并指出评估有效性高度依赖任务与提示设计。
原文摘要 · Abstract (English)
Large language models are increasingly being used in patient-facing medical question answering, where hallucinated outputs can vary widely in potential harm. However, existing hallucination standards and evaluation metrics focus primarily on factual correctness, treating all errors as equally severe. This obscures clinically relevant failure modes, particularly when models generate unsupported but actionable medical language. We propose a risk-sensitive evaluation framework that quantifies hallucinations through the presence of risk-bearing language, including treatment directives, contraindications, urgency cues, and mentions of high-risk medications. Rather than assessing clinical correctness, our approach evaluates the potential impact of hallucinated content if acted upon. We further combine risk scoring with a relevance measure to identify high-risk, low-grounding failures. We apply this framework to three instruction-tuned language models using controlled patient-facing prompts designed as safety stress tests. Our results show that models with similar surface-level behavior exhibit substantially different risk profiles and that standard evaluation metrics fail to capture these distinctions. These findings highlight the importance of incorporating risk sensitivity into hallucination evaluation and suggest that evaluation validity is critically dependent on task and prompt design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。