轻量化模型实现皮肤癌精准诊断,视觉与临床语义对齐更可信
SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis
- 冻结主干模型+低秩微调,大幅降低计算开销
- 在两个数据集上准确率超越大模型4.3%-6.2%,参数减少43%
- 生成可解释的视觉依据,提升医生信任度
将视觉语言模型应用于皮肤病学面临高计算成本、极端数据稀缺和深度学习黑箱三大难题。为此,我们提出SkinCLIP-VL——一种资源高效的框架,通过冻结感知与自适应推理范式,将冻结的CLIP编码器与轻量化的量化Qwen2.5-VL结合,采用低秩适配(LoRA)进行微调。为在长尾分布下严格对齐视觉区域与临床语义,提出一致性感知聚焦对齐(CFA)损失,该目标融合焦点重加权、语义对齐与校准机制。在ISIC和Derm7pt基准上,SkinCLIP-VL以43%更少参数量实现比130亿参数基线高出4.3%-6.2%的准确率。关键的是,盲测专家评估与分布外测试表明,其生成的视觉可解释理由显著提升了临床可信度,优于传统显著图。
原文摘要 · Abstract (English)
The deployment of vision-language models (VLMs) in dermatology is hindered by the trilemma of high computational costs, extreme data scarcity, and the black-box nature of deep learning. To address these challenges, we present SkinCLIP-VL, a resource-efficient framework that adapts foundation models for trustworthy skin cancer diagnosis. Adopting a frozen perception, adaptive reasoning paradigm, we integrate a frozen CLIP encoder with a lightweight, quantized Qwen2.5-VL via low-rank adaptation (LoRA). To strictly align visual regions with clinical semantics under long-tailed distributions, we propose the Consistency-aware Focal Alignment (CFA) Loss. This objective synergizes focal re-weighting, semantic alignment, and calibration. On ISIC and Derm7pt benchmarks, SkinCLIP-VL surpasses 13B-parameter baselines by 4.3-6.2% in accuracy with 43% fewer parameters. Crucially, blinded expert evaluation and out-of-distribution testing confirm that our visually grounded rationales significantly enhance clinical trust compared to traditional saliency maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。