arXiv:2605.30646cs.CLcs.AI2026-05被引 1

测试临床大模型对语义相同但表述不同的输入是否稳定,发现领域专用模型表现不一。

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

论文配图:Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs
图 1 · 摘自论文原文
  • 用自然语言推理技术筛选真正语义一致的改写提示
  • 16个模型在诊断数据集上测试,发现领域模型鲁棒性差异大
  • 适合关注医疗AI安全性的研究者和开发者参考

大型语言模型在临床应用中日益普及,但其行为对细微的语言变化(如换词、句式调整)高度敏感。这种敏感性在高风险医疗场景中构成隐患,因为语义等价的输入应产生一致预测。然而,关键挑战在于如何确保提示变化确实保留临床含义——现有基于嵌入的相似度度量常无法识别否定、时间或严重程度的差异。为此,我们提出一种基于自然语言推理(NLI)的语义验证框架,结合大模型判官与临床专家审核,筛选出真正语义不变的提示变体。同时引入三项指标量化模型敏感性:语义保持变体敏感度(MVS)、置信度变化(ΔC)和最坏情况不稳定性(WCI)。我们在相同模型族与参数规模下,评估了16个开源通用(GP)与医学专用(DS)模型,使用来自DiagnosisQA和MedQA数据集的重构提示。结果显示,不同领域模型间的鲁棒性表现混杂且高度依赖具体模型,即领域专业化并非始终提升或降低对语义保持提示重写下的鲁棒性;部分DS模型表现优异,而一些强基线通用模型也保持竞争力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior remains highly sensitive to subtle linguistic variations, such as rephrasing or syntactic variation. This sensitivity poses risks in safety-critical healthcare settings, where semantically equivalent inputs should produce consistent predictions. However, a key challenge is to ensure that prompt variations truly preserve clinical meaning, as embedding-based similarity metrics often fail to capture distinctions involving negation, temporality, or severity. To address this limitation, we propose a semantic verification framework based on Natural Language Inference (NLI) to filter meaning-preserving prompt variations, which are further refined using an LLM-as-a-judge and audited by a clinical expert. In addition, we introduce three metrics to quantify model sensitivity: MeaningPreserving Variation Sensitivity (MVS), confidence variation (ΔC), and Worst-Case Instability (WCI). We evaluate 16 open-source general-purpose (GP) and medical LLMs within the same model families and parameter scales, using reformulated prompts derived from the DiagnosisQA and MedQA datasets. Our results demonstrate that robustness differences between domain-specific (DS) models are mixed and highly model-dependent, i.e., domain specialization does not consistently improve or reduce robustness to meaning-preserving prompt reformulations. Several DS models rank among the most robust (when compared with GP counterparts), and strong GP baselines remain competitive as well.

医疗LLM语义稳定性模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。