无需反向传播,用可读词典实现医疗图像的透明化快速建模。
Toward Aristotelian Medical Representations: Backpropagation-Free Layer-wise Analysis for Interpretable Generalized Metric Learning on MedMNIST
- 基于预训练视觉模型构建通用表示空间,避免梯度微调
- 在MedMNIST v2上达到与基准相当的准确率
- 适合需要高透明度的临床医疗场景
尽管深度学习在医学影像中取得显著进展,但基于反向传播的模型仍存在“黑箱”问题,阻碍其临床应用。为此,我们提出阿瑞斯托泰利安快速对象建模(A-ROM),建立在柏拉图表征假说(PRH)基础上,该假说认为在大规模多样化数据上训练的模型会收敛到对现实的普遍客观表征。通过利用预训练视觉变换器(ViTs)的通用度量空间,A-ROM 实现了无需计算开销或模型不透明性的新医疗概念快速建模。我们以人类可读的概念词典和k近邻(kNN)分类器替代传统隐式决策层,确保模型逻辑可解释。在MedMNIST v2套件上的实验表明,A-ROM 在性能上可与标准基准竞争,同时提供简单、可扩展的少样本解决方案,满足现代临床环境对透明性的严格要求。
原文摘要 · Abstract (English)
While deep learning has achieved remarkable success in medical imaging, the "black-box" nature of backpropagation-based models remains a significant barrier to clinical adoption. To bridge this gap, we propose Aristotelian Rapid Object Modeling (A-ROM), a framework built upon the Platonic Representation Hypothesis (PRH). This hypothesis posits that models trained on vast, diverse datasets converge toward a universal and objective representation of reality. By leveraging the generalizable metric space of pretrained Vision Transformers (ViTs), A-ROM enables the rapid modeling of novel medical concepts without the computational burden or opacity of further gradient-based fine-tuning. We replace traditional, opaque decision layers with a human-readable concept dictionary and a k-Nearest Neighbors (kNN) classifier to ensure the model's logic remains interpretable. Experiments on the MedMNIST v2 suite demonstrate that A-ROM delivers performance competitive with standard benchmarks while providing a simple and scalable, "few-shot" solution that meets the rigorous transparency demands of modern clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。