arXiv:2601.12471cs.CLcs.AI2026-01Conference of the …被引 14

让医疗大模型学会在不确定时放弃回答,提升临床应用安全性

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

  • 设计统一评估框架,融合置信度校准与对抗扰动测试
  • 顶尖模型在不确定时仍常错误作答,显式弃权选项显著提升安全率
  • 适合关注医疗AI安全、模型可信决策的研究者和开发者

当前大型语言模型(LLMs)的评估主要关注准确性,但在真实且高风险的临床场景中,模型在不确定时主动放弃回答同样关键。本文提出 MedAbstain,一个统一的基准与评估协议,用于医学多选题问答(MCQA)中的弃权行为——该任务可泛化至代理式决策选择。该框架整合了置信度校准、对抗性问题扰动和显式弃权选项。对开源与闭源模型的系统评估显示,即使最先进的高精度模型也常在不确定时强行作答。值得注意的是,提供显式弃权选项能显著提高模型不确定性感知并促进更安全的弃权行为,远优于输入扰动;而模型规模扩大或高级提示策略带来的改进有限。研究强调了弃权机制在可信部署中的核心作用,并为高风险应用场景提供了实用安全优化路径。

原文摘要 · Abstract (English)

Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.

医疗AI模型安全弃权机制可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。