arXiv:2508.20188cs.CVcs.LG2025-08

让AI皮肤诊断更可信:用量化特征提升大模型可解释性

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

  • 用图像预测皮肤病变的量化属性(如面积),微调多模态大模型
  • 在SLICE-3D数据集上,基于属性的检索准确率显著提升
  • 适合临床辅助诊断与AI可解释性研究者

人工智能在皮肤疾病诊断(包括癌症)中已展现显著成效,具备辅助临床分析的潜力。然而,模型预测的可解释性仍需大幅提升才能实际应用。为此,本文探索结合多模态大语言模型(MLLMs)与定量属性分析的路径。MLLMs可通过自然语言提供诊断推理过程,增强可解释性;而近期研究发现,病变面积等外观量化属性对恶性程度预测具有高准确性。将模型预测基于这些概念,有望提升可解释性。本文证明,通过微调使MLLM从图像中预测定量属性值,可实现其嵌入空间的属性对齐。我们以SLICE-3D数据集为基础,开展属性特定的内容检索案例研究,验证了该方法的有效性。

原文摘要 · Abstract (English)

Artificial Intelligence models have demonstrated significant success in diagnosing skin diseases, including cancer, showing the potential to assist clinicians in their analysis. However, the interpretability of model predictions must be significantly improved before they can be used in practice. To this end, we explore the combination of two promising approaches: Multimodal Large Language Models (MLLMs) and quantitative attribute usage. MLLMs offer a potential avenue for increased interpretability, providing reasoning for diagnosis in natural language through an interactive format. Separately, a number of quantitative attributes that are related to lesion appearance (e.g., lesion area) have recently been found predictive of malignancy with high accuracy. Predictions grounded as a function of such concepts have the potential for improved interpretability. We provide evidence that MLLM embedding spaces can be grounded in such attributes, through fine-tuning to predict their values from images. Concretely, we evaluate this grounding in the embedding space through an attribute-specific content-based image retrieval case study using the SLICE-3D dataset.

皮肤诊断多模态大模型可解释性量化属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。