构建皮肤病罕见病诊断评估基准,揭示大模型推理短板
Mind the Rarities: Can Rare Skin Diseases Be Reliably Diagnosed via Diagnostic Reasoning?
- 基于真实病例构建多模态长文本数据集,标注诊断推理链
- 22个主流大模型在罕见病诊断中准确率与推理质量普遍偏低
- 提出医生对齐的评估指标,适合临床辅助诊断研究者使用
大型视觉语言模型在皮肤科表现出色,但针对罕见病的诊断推理评估仍处于空白。现有基准集中于常见疾病,仅评估最终准确率,忽略临床推理过程,而该过程对复杂病例至关重要。本文构建了DermCase,一个源自同行评审病例报告的长上下文基准,包含26,030个多模态图像-文本对和6,354个临床挑战性病例,每例均标注完整临床信息及逐步推理链条。为实现可靠评估,我们建立基于DermLIP的相似性度量,其对差异诊断质量的评估与皮肤科医生高度一致。对22个领先的大模型进行基准测试,发现其在诊断准确率、差异诊断和临床推理方面均存在显著缺陷。微调实验表明,指令微调显著提升性能,而直接偏好优化(DPO)收益甚微。系统性错误分析进一步揭示当前模型推理能力的关键局限。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) demonstrate strong performance in dermatology; however, evaluating diagnostic reasoning for rare conditions remains largely unexplored. Existing benchmarks focus on common diseases and assess only final accuracy, overlooking the clinical reasoning process, which is critical for complex cases. We address this gap by constructing DermCase, a long-context benchmark derived from peer-reviewed case reports. Our dataset contains 26,030 multi-modal image-text pairs and 6,354 clinically challenging cases, each annotated with comprehensive clinical information and step-by-step reasoning chains. To enable reliable evaluation, we establish DermLIP-based similarity metrics that achieve stronger alignment with dermatologists for assessing differential diagnosis quality. Benchmarking 22 leading LVLMs exposes significant deficiencies across diagnosis accuracy, differential diagnosis, and clinical reasoning. Fine-tuning experiments demonstrate that instruction tuning substantially improves performance while Direct Preference Optimization (DPO) yields minimal gains. Systematic error analysis further reveals critical limitations in current models' reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。