测试大模型在模糊病患自述下的诊断能力,发现真实语境下表现显著下降。
From Fuzzy Speech to Medical Insight: Benchmarking LLMs on Noisy Patient Narratives
- 构建合成噪声数据集,模拟患者自述中的模糊语言和非专业术语
- 在噪声语境下,主流大模型诊断准确率下降超30%
- 适合研究医疗AI鲁棒性或临床场景落地的开发者使用
大型语言模型(LLMs)在医疗领域的广泛应用引发对其解读患者自述能力的关切。现有基准多基于清晰、结构化的临床文本,难以反映真实场景。本文提出一种新型合成数据集,模拟患者自述中不同水平的语言噪声、模糊表达及非专业术语。数据集包含具有真实临床一致性的病例,标注有真实诊断结果,覆盖从清晰到模糊的多种沟通风格。我们在此基准上微调并评估多个前沿模型(包括BERT与T5类编码器-解码器模型)。为支持可复现研究,我们发布「噪声诊断基准」(Noisy Diagnostic Benchmark, NDB),一个结构化合成数据集,用于在真实语言条件下压力测试和比较大模型的诊断能力。该基准已开源:https://github.com/lielsheri/PatientSignal。
原文摘要 · Abstract (English)
The widespread adoption of large language models (LLMs) in healthcare raises critical questions about their ability to interpret patient-generated narratives, which are often informal, ambiguous, and noisy. Existing benchmarks typically rely on clean, structured clinical text, offering limited insight into model performance under realistic conditions. In this work, we present a novel synthetic dataset designed to simulate patient self-descriptions characterized by varying levels of linguistic noise, fuzzy language, and layperson terminology. Our dataset comprises clinically consistent scenarios annotated with ground-truth diagnoses, spanning a spectrum of communication clarity to reflect diverse real-world reporting styles. Using this benchmark, we fine-tune and evaluate several state-of-the-art models (LLMs), including BERT-based and encoder-decoder T5 models. To support reproducibility and future research, we release the Noisy Diagnostic Benchmark (NDB), a structured dataset of noisy, synthetic patient descriptions designed to stress-test and compare the diagnostic capabilities of large language models (LLMs) under realistic linguistic conditions. We made the benchmark available for the community: https://github.com/lielsheri/PatientSignal
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。