为皮肤科多模态大模型设计可信评估体系,实现精准诊断文本评测。
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- 构建4000张真实皮肤图像与专家认证诊断文本的基准数据集
- 自动评估器误差仅0.117(满分5分),接近专家水平
- 支持逐病例细粒度分析,适合临床部署前模型验证
多模态大语言模型正被用于直接从图像生成皮肤科诊断文本,但可靠评估仍是临床应用的主要瓶颈。本文提出结合DermBench(精心构建的基准)与DermEval(鲁棒的自动评估器)的新框架,实现可临床解读、可复现且可扩展的评估。DermBench包含4,000张真实皮肤图像与专家认证诊断文本,并使用基于LLM的评判系统,在临床相关维度上对候选文本评分,确保评估一致性与全面性。针对单例评估,训练了DermEval——一个无参考的多模态评估器,给定图像和生成文本后,输出结构化批评及综合得分与各维度评分。该能力支持细粒度病例分析,有助于发现模型缺陷与偏见。在包含4,500例的多样化数据集上实验表明,DermBench与DermEval与专家评分的平均偏差分别为0.251和0.117(满分5分),能可靠衡量不同多模态大模型的诊断能力与可信度。
原文摘要 · Abstract (English)
Multimodal large language models (LLMs) are increasingly used to generate dermatology diagnostic narratives directly from images. However, reliable evaluation remains the primary bottleneck for responsible clinical deployment. We introduce a novel evaluation framework that combines DermBench, a meticulously curated benchmark, with DermEval, a robust automatic evaluator, to enable clinically meaningful, reproducible, and scalable assessment. We build DermBench, which pairs 4,000 real-world dermatology images with expert-certified diagnostic narratives and uses an LLM-based judge to score candidate narratives across clinically grounded dimensions, enabling consistent and comprehensive evaluation of multimodal models. For individual case assessment, we train DermEval, a reference-free multimodal evaluator. Given an image and a generated narrative, DermEval produces a structured critique along with an overall score and per-dimension ratings. This capability enables fine-grained, per-case analysis, which is critical for identifying model limitations and biases. Experiments on a diverse dataset of 4,500 cases demonstrate that DermBench and DermEval achieve close alignment with expert ratings, with mean deviations of 0.251 and 0.117 (out of 5), respectively, providing reliable measurement of diagnostic ability and trustworthiness across different multimodal LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。