构建患者图文问答与皮肤病变分割数据集,推动精准皮肤病诊疗研究
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
- 设计临床标准评估框架DAS,结构化提取36项高阶与27项细粒度皮肤特征
- 提出多模态模型在问答任务中准确率达79.8%,分割任务中骰子系数达0.566
- 公开数据集与评估协议,助力面向患者的皮肤病视觉语言模型发展
近期皮肤病图像分析进展依赖大规模标注数据集,但现有基准多聚焦于皮肤镜图像,缺乏患者自述问题和临床背景,限制其在以患者为中心的医疗中的应用。为此,我们推出DermaVQA-DAS,扩展自DermaVQA数据集,支持闭合式问答(QA)与皮肤病变分割两项互补任务。核心是专家制定的皮肤病评估框架DAS,系统化、标准化地捕捉临床相关的皮肤特征,包含36个高阶与27个细粒度评估问题,提供中英文多选选项。基于DAS,我们构建了专家标注的问答与分割数据集,并对主流多模态模型进行基准测试。分割任务中,评估多种提示策略,发现提示设计影响性能:默认提示在均值最大与均值平均聚合方案下表现最佳;结合患者查询标题与内容的增强提示在多数投票微评分下最优,使用BiomedParse模型达到0.395的杰卡德指数与0.566的骰子系数。闭合式问答任务整体表现良好,各模型平均准确率介于0.729至0.798之间,o3表现最佳(0.798),紧随其后为GPT-4.1(0.796),Gemini-1.5-Pro在同系列中表现优异(0.783)。我们公开发布DermaVQA-DAS、DAS框架及评估协议,以推动患者中心型皮肤病视觉语言建模研究(https://osf.io/72rp3)。
原文摘要 · Abstract (English)
Recent advances in dermatological image analysis have been driven by large-scale annotated datasets; however, most existing benchmarks focus on dermatoscopic images and lack patient-authored queries and clinical context, limiting their applicability to patient-centered care. To address this gap, we introduce DermaVQA-DAS, an extension of the DermaVQA dataset that supports two complementary tasks: closed-ended question answering (QA) and dermatological lesion segmentation. Central to this work is the Dermatology Assessment Schema (DAS), a novel expert-developed framework that systematically captures clinically meaningful dermatological features in a structured and standardized form. DAS comprises 36 high-level and 27 fine-grained assessment questions, with multiple-choice options in English and Chinese. Leveraging DAS, we provide expert-annotated datasets for both closed QA and segmentation and benchmark state-of-the-art multimodal models. For segmentation, we evaluate multiple prompting strategies and show that prompt design impacts performance: the default prompt achieves the best results under Mean-of-Max and Mean-of-Mean evaluation aggregation schemes, while an augmented prompt incorporating both patient query title and content yields the highest performance under majority-vote-based microscore evaluation, achieving a Jaccard index of 0.395 and a Dice score of 0.566 with BiomedParse. For closed-ended QA, overall performance is strong across models, with average accuracies ranging from 0.729 to 0.798; o3 achieves the best overall accuracy (0.798), closely followed by GPT-4.1 (0.796), while Gemini-1.5-Pro shows competitive performance within the Gemini family (0.783). We publicly release DermaVQA-DAS, the DAS schema, and evaluation protocols to support and accelerate future research in patient-centered dermatological vision-language modeling (https://osf.io/72rp3).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。