用医学教材指导推理,提升皮肤科诊断模型准确性和可信度。
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
- 基于权威教材生成分层诊断路径,提供临床可解释的推理监督。
- 在多个皮肤科数据集上显著优于现有医疗视觉语言模型。
- 适合需要可解释诊断结果的研究者和临床辅助系统开发者。
视觉-语言模型(VLMs)在皮肤科诊断辅助中展现出潜力,但其可信度与临床实用性受限于三大挑战:数据集异质性导致标签与概念标注不一致、缺乏可验证的诊断推理依据、小规模密集标注数据向大规模稀疏标注数据迁移时扩展性差。为此,我们提出Skin-R1,一种面向皮肤科的VLM,通过融合教材引导的临床推理监督与强化学习(RL),提升诊断预测的准确性与鲁棒性。首先,构建基于教材的推理生成器,合成具有层次结构和鉴别诊断(DDx)特征的诊断路径;其次,利用这些路径进行监督微调(SFT),建立临床可解释的推理基础;最后,设计融入疾病层次结构的强化学习奖励机制,使模型能将可解释推理推广至大规模稀疏标注数据。多组实验表明,Skin-R1在多个皮肤科基准测试中均显著优于当前最优医学视觉语言模型。消融研究进一步验证了SFT阶段引入的基于知识的推理监督的关键作用。
原文摘要 · Abstract (English)
Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis. However, their trustworthiness and clinical utility remain limited by three key challenges: heterogeneous datasets with inconsistent diagnostic labels and concept annotations, the lack of grounded diagnostic rationales for reliable reasoning supervision, and limited scalability when transferring knowledge from small, densely annotated datasets to large collections with sparse labels. To address these challenges, we propose Skin-R1, a dermatology-oriented VLM that integrates textbook-grounded clinical reasoning supervision with reinforcement learning (RL) to improve the accuracy and robustness of diagnostic prediction. First, we construct a textbook-based reasoning generator that synthesizes hierarchy-aware and differential-diagnosis (DDx) diagnostic trajectories derived from authoritative dermatology knowledge. Second, these trajectories are used for supervised fine-tuning (SFT), establishing a clinically grounded reasoning foundation for the model. Finally, we introduce an RL training framework that incorporates the hierarchical structure of dermatological diseases into the reward design, enabling the model to generalize grounded diagnostic reasoning to large-scale datasets with sparse annotations. Extensive experiments across multiple dermatology benchmarks demonstrate that Skin-R1 consistently improves diagnostic accuracy and robustness compared to state-of-the-art Med-VLM baselines. Ablation studies further highlight the critical role of grounded reasoning supervision introduced during the SFT stage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。