arXiv:2601.09853cs.CLcs.AI2026-01ACL

测试大模型在真实医疗提问中的纠错能力,发现其常忽略错误前提导致误导。

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

  • 构建半自动化数据集MedRedFlag,收录1100+需纠正误解的患者提问
  • 大模型在检测到错误前提后仍常直接回答原问题,而非引导纠正
  • 结果揭示医疗AI在真实沟通中存在严重安全缺陷,适合关注医疗AI风险的研究者

患者在现实健康咨询中常隐含错误假设。安全的医疗沟通应先纠正误解,再回应真实需求。尽管大语言模型(LLMs)被广泛用于获取医疗建议,但其在处理此类问题上的能力尚未验证。本文构建了MedRedFlag数据集,包含1100+条来自Reddit的真实健康提问,均需通过引导式回应纠正错误前提。我们系统比较了先进LLMs与临床医生的响应表现。分析发现,即使模型识别出错误前提,仍频繁直接作答,可能导致次优医疗决策。该基准揭示了当前医疗类LLMs在真实语境下的显著能力缺口,凸显面向患者的医疗AI系统存在关键安全隐患。代码与数据集已开源。

原文摘要 · Abstract (English)

Real-world health questions from patients often unintentionally embed false assumptions or premises. In such cases, safe medical communication typically involves redirection: addressing the implicit misconception and then responding to the underlying patient context, rather than the original question. While large language models (LLMs) are increasingly being used by lay users for medical advice, they have not yet been tested for this crucial competency. Therefore, in this work, we investigate how LLMs react to false premises embedded within real-world health questions. We develop a semi-automated pipeline to curate MedRedFlag, a dataset of 1100+ questions sourced from Reddit that require redirection. We then systematically compare responses from state-of-the-art LLMs to those from clinicians. Our analysis reveals that LLMs often fail to redirect problematic questions, even when the problematic premise is detected, and provide answers that could lead to suboptimal medical decision making. Our benchmark and results reveal a novel and substantial gap in how LLMs perform under the conditions of real-world health communication, highlighting critical safety concerns for patient-facing medical AI systems. Code and dataset are available at https://github.com/srsambara-1/MedRedFlag.

医疗AI大模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。