arXiv:2507.08030cs.CLcs.CE2025-07被引 9

AI医学模型输出的警示语逐年减少,2025年多数已无警示。

A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models

  • 对比2022至2025年模型输出,用500张乳腺影像等数据检测警示语
  • 大语言模型警示语从26.3%降至0.97%,视觉语言模型从19.6%降至1.05%
  • 模型越强大越需动态警示,避免误导临床决策

生成式AI模型(包括大语言模型LLMs和视觉语言模型VLMs)在医学影像解读和临床问答中的应用日益广泛。由于其输出常含错误,医疗警示语至关重要,可提醒用户AI结果未经专业审核且不可替代医疗建议。本研究评估了2022至2025年间多代模型在医学图像与问题上的警示语存在情况。基于500张乳腺钼靶、500张胸片、500张皮肤科图像及500个医学问题,分析输出中是否包含警示语。结果显示,LLM的警示语占比从2022年的26.3%下降至2025年的0.97%,VLM则从2023年的19.6%降至2025年的1.05%。至2025年,多数模型已不再显示警示语。随着公开模型能力提升与权威感增强,警示机制必须根据具体临床场景动态调整,以保障医疗安全。

原文摘要 · Abstract (English)

Generative AI models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used to interpret medical images and answer clinical questions. Their responses often include inaccuracies; therefore, safety measures like medical disclaimers are critical to remind users that AI outputs are not professionally vetted or a substitute for medical advice. This study evaluated the presence of disclaimers in LLM and VLM outputs across model generations from 2022 to 2025. Using 500 mammograms, 500 chest X-rays, 500 dermatology images, and 500 medical questions, outputs were screened for disclaimer phrases. Medical disclaimer presence in LLM and VLM outputs dropped from 26.3% in 2022 to 0.97% in 2025, and from 19.6% in 2023 to 1.05% in 2025, respectively. By 2025, the majority of models displayed no disclaimers. As public models become more capable and authoritative, disclaimers must be implemented as a safeguard adapting to the clinical context of each output.

AI医疗安全警示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。