arXiv:2601.14084cs.CVcs.AI2026-01被引 2

面向皮肤科视觉问答的临床标注数据集,助力评估多模态模型的医学推理能力。

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

  • 由皮肤科医生标注,覆盖570名患者、656张图像的多层次问题与答案。
  • 包含14.474条问答式标注,涵盖诊断、形态、分布等22个核心临床问题。
  • 适合研究医疗视觉理解、临床推理的AI开发者和医学人工智能团队。

视觉语言模型在医疗领域日益重要,但其在皮肤科领域的评估仍受限于以图像级分类(如病灶识别)为主的数据库。此类数据无法评估多模态模型在视觉理解、语言对齐和临床推理方面的综合能力。为此,我们提出DermaBench,一个基于Diverse Dermatology Images(DDI)数据集构建的皮肤科视觉问答基准。该数据集包含656张来自570名独特患者的临床图像,覆盖Fitzpatrick皮肤类型I-VI。通过22个主问题(单选、多选、开放式)的分层标注体系,专家皮肤科医生为每张图像标注了诊断、解剖部位、病灶形态、分布、表面特征、颜色及图像质量,并提供开放式叙述描述与摘要,共生成约14.474条VQA风格标注。为尊重上游许可,DermaBench以元数据形式发布,公开获取于Harvard Dataverse。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.

皮肤科视觉问答临床标注多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。