arXiv:2504.01201cs.CLcs.AI2025-04被引 20

医学大模型易受无关信息干扰,实测准确率下降超17%。

Medical large language models are easily distracted

  • 构建医疗干扰问答基准MedDistractQA,模拟真实临床噪声
  • 干扰词使模型准确率最高下降17.9%,现有优化方法无效甚至恶化
  • 揭示模型缺乏逻辑筛选能力,适用于临床应用研究者

大型语言模型(LLM)有潜力变革医疗,但真实临床场景包含大量无关信息,可能影响其表现。随着环境语音转录等辅助技术的发展,从患者对话中自动生成草稿病历,可能引入额外噪声,因此评估LLM过滤相关数据的能力至关重要。为此,我们开发了MedDistractQA基准,采用类似USMLE的题目并嵌入模拟真实世界的干扰信息。研究发现,干扰性陈述(如具有临床含义的多义词用于非临床语境,或提及无关健康状况)可使LLM准确率最高降低17.9%。常见的性能提升方法,如检索增强生成(RAG)和医学微调,并未缓解此问题,反而在某些情况下引入新混杂因素,进一步降低性能。结果表明,LLM天生缺乏区分相关与无关临床信息的逻辑机制,对实际应用构成挑战。MedDistractQA及研究成果凸显了提升LLM对冗余信息鲁棒性的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) have the potential to transform medicine, but real-world clinical scenarios contain extraneous information that can hinder performance. The rise of assistive technologies like ambient dictation, which automatically generates draft notes from live patient encounters, has the potential to introduce additional noise making it crucial to assess the ability of LLM's to filter relevant data. To investigate this, we developed MedDistractQA, a benchmark using USMLE-style questions embedded with simulated real-world distractions. Our findings show that distracting statements (polysemous words with clinical meanings used in a non-clinical context or references to unrelated health conditions) can reduce LLM accuracy by up to 17.9%. Commonly proposed solutions to improve model performance such as retrieval-augmented generation (RAG) and medical fine-tuning did not change this effect and in some cases introduced their own confounders and further degraded performance. Our findings suggest that LLMs natively lack the logical mechanisms necessary to distinguish relevant from irrelevant clinical information, posing challenges for real-world applications. MedDistractQA and our results highlights the need for robust mitigation strategies to enhance LLM resilience to extraneous information.

大模型医疗AI抗干扰评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。