HierViT用原型可视化让医学图像分类模型像人一样思考并解释决策。
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification
- 构建分层结构,用人类定义的特征原型进行推理。
- 在两个医学数据集上达到先进准确率,且解释与医生判断一致。
- 通过原型和注意力热图双重机制,实现可解释性与性能兼顾。
可解释性是医疗等高风险领域的重要需求。视觉变压器主要依赖注意力提取来揭示模型推理过程。本文提出HierViT,一种兼具高性能与新可解释能力的视觉变压器。其采用分层结构处理领域特异性特征以实现预测,从设计上具备可解释性:通过人类定义的特征原型(以示例图像呈现)推导目标输出。结合领域知识使模型推理在语义上贴近人类思维,因而直观易懂。同时,注意力热图可视化识别各特征的关键区域,为验证预测提供多功能工具。在两个医学基准数据集LIDC-IDRI(肺结节评估)和derm7pt(皮肤病变分类)上,HierViT分别实现了优于或相当的预测精度,且解释结果与人类认知相符。
原文摘要 · Abstract (English)
Explainability is a highly demanded requirement for applications in high-risk areas such as medicine. Vision Transformers have mainly been limited to attention extraction to provide insight into the model's reasoning. Our approach combines the high performance of Vision Transformers with the introduction of new explainability capabilities. We present HierViT, a Vision Transformer that is inherently interpretable and adapts its reasoning to that of humans. A hierarchical structure is used to process domain-specific features for prediction. It is interpretable by design, as it derives the target output with human-defined features that are visualized by exemplary images (prototypes). By incorporating domain knowledge about these decisive features, the reasoning is semantically similar to human reasoning and therefore intuitive. Moreover, attention heatmaps visualize the crucial regions for identifying each feature, thereby providing HierViT with a versatile tool for validating predictions. Evaluated on two medical benchmark datasets, LIDC-IDRI for lung nodule assessment and derm7pt for skin lesion classification, HierViT achieves superior and comparable prediction accuracy, respectively, while offering explanations that align with human reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。