用大模型评估皮肤病模型的解释是否准确聚焦病灶区
LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models
- 结合颜色显著性与面部结构先验生成可临床解读的注意力图
- 通过大模型评判,发现新方法在病灶定位上更可信
- 适合医学AI可解释性研究者参考
本研究提出一种面向面部皮肤病诊断的领域专用大语言模型(LLM)视觉解释评估框架,用于评估视觉注意力解释的可靠性。现有研究多关注分类性能提升,较少系统检验解释是否对应临床相关病灶区域。本文开发了一种图像驱动的注意力生成算法,从面部皮肤疾病图像中生成聚焦病灶的注意力图。该方法融合颜色显著性、面部空间先验、高斯平滑与注意力叠加可视化,优于传统的基于梯度的解释方法。此外,设计了基于GPT-5.5、Gemini 3.5 Flash和Claude Sonnet 4.6的LLM-as-a-Judge评估框架,从病灶定位准确性和解释可信度两个角度评估生成的解释。
原文摘要 · Abstract (English)
This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing visual attention explanations in facial skin disease diagnosis. While previous studies have primarily focused on improving classification performance, relatively few studies have systematically examined whether visual explanations are grounded in clinically relevant lesion regions. In this study, an image-driven visual attention generation algorithm was developed to produce lesion-focused attention maps from facial skin disease images. Unlike conventional gradient-based explainability methods, the proposed approach combines color saliency, facial spatial priors, Gaussian smoothing, and attention overlay visualization to generate clinically interpretable attention maps. Furthermore, an LLM-as-a-Judge evaluation framework was designed using GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6 to assess the generated visual explanations from the perspectives of lesion localization and explanation trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。