arXiv:2511.10912cs.CLcs.AI2025-11被引 1

用美剧《豪斯医生》数据评估大模型罕见病诊断能力,发现新模型性能提升2.3倍。

Evaluating Large Language Models on Rare Disease Diagnosis: A Case Study using House M.D

  • 基于美剧情节构建176例罕见病诊断数据集,用于测试模型叙事推理能力。
  • 模型准确率在16.48%至38.64%之间,最新一代模型比旧版提升2.3倍。
  • 为医学AI研究提供可复现的评测框架,适合医疗AI与临床辅助系统开发者参考。

大型语言模型(LLMs)在多个领域展现出强大能力,但在从叙述性医学案例中诊断罕见疾病方面仍缺乏充分研究。本文引入一个新颖的数据集,包含从医疗电视剧《豪斯医生》中提取的176个症状-诊断对,该数据集已在医学教育中被验证可用于罕见病识别教学。我们评估了四种前沿LLM(如GPT 4o mini、GPT 5 mini、Gemini 2.5 Flash和Gemini 2.5 Pro)在基于叙述的诊断推理任务上的表现。结果显示性能差异显著,准确率介于16.48%至38.64%之间,较旧版本模型的新一代架构实现2.3倍性能提升。尽管所有模型在罕见病诊断上仍面临重大挑战,但跨架构的改进趋势表明未来发展方向。本研究提供的教育验证基准,建立了叙述性医学推理的基线性能指标,并提供了公开可访问的评估框架,以推动人工智能辅助诊断研究的发展。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated capabilities across diverse domains, yet their performance on rare disease diagnosis from narrative medical cases remains underexplored. We introduce a novel dataset of 176 symptom-diagnosis pairs extracted from House M.D., a medical television series validated for teaching rare disease recognition in medical education. We evaluate four state-of-the-art LLMs such as GPT 4o mini, GPT 5 mini, Gemini 2.5 Flash, and Gemini 2.5 Pro on narrative-based diagnostic reasoning tasks. Results show significant variation in performance, ranging from 16.48% to 38.64% accuracy, with newer model generations demonstrating a 2.3 times improvement. While all models face substantial challenges with rare disease diagnosis, the observed improvement across architectures suggests promising directions for future development. Our educationally validated benchmark establishes baseline performance metrics for narrative medical reasoning and provides a publicly accessible evaluation framework for advancing AI-assisted diagnosis research.

罕见病诊断大模型评测医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。