通过真实临床场景的扰动数据,揭示大模型与医生在诊断决策上的差异。
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
- 构建三类扰动的临床病例数据集,模拟性别、语风、格式变化。
- 模型对性别和语风更敏感,医生对模型生成的摘要更敏感。
- 适合医疗AI评估、临床决策研究者参考。
临床鲁棒性对医疗大模型的安全部署至关重要,但关于模型与人类在真实临床变异性下的响应差异仍存疑问。为此,我们提出MedPerturb数据集,系统评估医疗LLMs在受控扰动下的表现。该数据集包含800个临床案例,覆盖多种病理,每例在三个维度上进行扰动:(1) 性别修改(如性别互换或移除);(2) 风格变化(如不确定表达或口语化语气);(3) 格式变更(如模型生成的多轮对话或摘要)。我们提供四款模型输出及每例三名专家人工标注。通过两个案例研究发现,模型对性别与风格扰动更敏感,而人类对模型生成的摘要等格式变化更敏感。结果表明,需建立超越静态基准的评估框架,以衡量模型与临床医生在真实临床变异性下的决策一致性。
原文摘要 · Abstract (English)
Clinical robustness is critical to the safe deployment of medical Large Language Models (LLMs), but key questions remain about how LLMs and humans may differ in response to the real-world variability typified by clinical settings. To address this, we introduce MedPerturb, a dataset designed to systematically evaluate medical LLMs under controlled perturbations of clinical input. MedPerturb consists of clinical vignettes spanning a range of pathologies, each transformed along three axes: (1) gender modifications (e.g., gender-swapping or gender-removal); (2) style variation (e.g., uncertain phrasing or colloquial tone); and (3) format changes (e.g., LLM-generated multi-turn conversations or summaries). With MedPerturb, we release a dataset of 800 clinical contexts grounded in realistic input variability, outputs from four LLMs, and three human expert reads per clinical context. We use MedPerturb in two case studies to reveal how shifts in gender identity cues, language style, or format reflect diverging treatment selections between humans and LLMs. We find that LLMs are more sensitive to gender and style perturbations while human annotators are more sensitive to LLM-generated format perturbations such as clinical summaries. Our results highlight the need for evaluation frameworks that go beyond static benchmarks to assess the similarity between human clinician and LLM decisions under the variability characteristic of clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。