arXiv:2604.10441cs.AI2026-04

用真实患者噪声测试医疗大模型,发现小模型更易崩塌。

VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise

论文配图:VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise
图 1 · 摘自论文原文
  • 构建可配置噪声框架,模拟患者记忆差、理解力弱等临床真实问题。
  • 所有模型在真实噪声下诊断准确率下降15%-25%,对话变长34%-55%。
  • 适合评估医疗AI鲁棒性,尤其关注小模型和真实临床场景的差距。

医疗大语言模型在标准评测中表现优异,但无法反映真实临床中患者存在的记忆缺失、健康素养不足、焦虑等沟通障碍。我们提出VeriSim,一个保持医学真值的患者仿真框架,通过混合UMLS与大模型验证机制,在患者回答中注入可控且具临床依据的噪声。该框架基于同行评审文献定义六类噪声维度,涵盖患者回忆能力、健康认知水平及因病耻感导致的隐瞒等问题。在七种开源大模型上的实验表明,所有模型在真实患者噪声下性能显著下降:诊断准确率降低15%-25%,对话长度增加34%-55%。小型模型(7B)的性能退化比大型模型(70B+)高出40%;在标准语料上进行医学微调对抵抗沟通噪声作用有限。经认证临床医生评估,仿真质量高,评分者间一致性良好(kappa > 0.80);LLM作为裁判也展现出可量化的可靠性,适用于大规模评估。结果揭示当前医疗AI存在显著的仿真到现实差距。我们开源VeriSim,为临床鲁棒性评估提供严格基准。

原文摘要 · Abstract (English)

Medical large language models (LLMs) achieve impressive performance on standardized benchmarks, yet these evaluations fail to capture the complexity of real clinical encounters where patients exhibit memory gaps, limited health literacy, anxiety, and other communication barriers. We introduce VeriSim, a truth-preserving patient simulation framework that injects controllable, clinically evidence-grounded noise into patient responses while maintaining strict adherence to medical ground truth through a hybrid UMLS-LLM verification mechanism. Our framework operationalizes six noise dimensions derived from peer-reviewed medical communication literature, capturing authentic clinical phenomena such as patient recall limitations, health literacy barriers, and stigma-driven non-disclosure. Experiments across seven open-weight LLMs reveal that all models degrade significantly under realistic patient noise, with diagnostic accuracy dropping 15-25% and conversation length increasing 34-55%. Notably, smaller models (7B) show 40% greater degradation than larger models (70B+), while medical fine-tuning on standard corpora provides limited robustness benefits against patient communication noise. Evaluation by board-certified clinicians demonstrates high-quality simulation with strong inter-annotator agreement (kappa > 0.80), while LLM-as-a-Judge serves as a validated auxiliary evaluator achieving comparable reliability for scalable assessment. Our results highlight a critical Sim-to-Real gap in current medical AI. We release VeriSim as an open-source noise-injection framework, establishing a rigorous testbed for evaluating clinical robustness.

医疗AI噪声测试大模型评估临床仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。