arXiv:2508.18430cs.CVcs.AI2025-08被引 1

用专家-通用框架提升皮肤科视觉问答的准确率与效率

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering

  • 专家模型快速诊断,通用模型生成解释,协同推理更精准
  • 诊断准确率比最强基线高18%,显存占用降低20%以上
  • 适合临床部署,兼顾精度、速度与可解释性

视觉语言模型在医疗任务中展现巨大潜力,但其通用特性限制了专业诊断精度,且模型过大导致实际临床部署成本高。为此,我们提出CLARIFY——一种用于皮肤科视觉问答(VQA)的专家-通用框架。该框架包含两个组件:(i) 轻量级、领域训练的图像分类器(专家),提供快速高精度诊断预测;(ii) 强大但压缩的对话式视觉语言模型(通用),生成自然语言解释。专家的预测直接引导通用模型推理,聚焦正确诊断路径。该协同机制通过基于知识图谱的检索模块进一步强化,使通用模型的回答基于真实皮肤病学知识,确保准确性和可靠性。这种分层设计不仅减少诊断错误,还显著提升计算效率。在自建多模态皮肤科数据集上的实验表明,CLARIFY相比最强基线(微调的未压缩单模型VLM)诊断准确率提升18%,平均显存需求降低至少20%,延迟减少5%以上。结果表明,专家-通用系统为构建轻量化、可信且临床可行的AI系统提供了有效范式。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown significant potential for medical tasks; however, their general-purpose nature can limit specialized diagnostic accuracy, and their large size poses substantial inference costs for real-world clinical deployment. To address these challenges, we introduce CLARIFY, a Specialist-Generalist framework for dermatological visual question answering (VQA). CLARIFY combines two components: (i) a lightweight, domain-trained image classifier (the Specialist) that provides fast and highly accurate diagnostic predictions, and (ii) a powerful yet compressed conversational VLM (the Generalist) that generates natural language explanations to user queries. In our framework, the Specialist's predictions directly guide the Generalist's reasoning, focusing it on the correct diagnostic path. This synergy is further enhanced by a knowledge graph-based retrieval module, which grounds the Generalist's responses in factual dermatological knowledge, ensuring both accuracy and reliability. This hierarchical design not only reduces diagnostic errors but also significantly improves computational efficiency. Experiments on our curated multimodal dermatology dataset demonstrate that CLARIFY achieves an 18\% improvement in diagnostic accuracy over the strongest baseline, a fine-tuned, uncompressed single-line VLM, while reducing the average VRAM requirement and latency by at least 20\% and 5\%, respectively. These results indicate that a Specialist-Generalist system provides a practical and powerful paradigm for building lightweight, trustworthy, and clinically viable AI systems.

皮肤科AI视觉问答专家-通用轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。