无需重训练,通过几何优化提升多模态医学影像异常检测效果
Towards Modality-Agnostic Medical Image Anomaly Detection: A Training-Free Manifold Refinement Approach
- 在特征提取后加入无参流形精炼阶段,压缩正常样本聚集区
- 跨五种模态七数据集测试,四项指标达最优,平均精度领先
- 可适配任意预训练模型,适合临床中无法重训的场景
将基于AI的异常检测应用于多样化的临床影像场景仍具挑战性,因多数方法依赖特定模态架构、解剖先验或大量重训练,难以作为通用筛查工具。单类分类(OCC)通过仅使用正常数据训练实现标签高效,但传统两阶段流程直接在预训练嵌入上拟合密度估计,未能利用潜在空间中的判别结构。本文提出一种无需训练、模态无关的框架,在特征提取与异常评分间引入显式流形精炼阶段。通过基于UMAP的邻域图估算的经验密度权重,迭代引导嵌入向局部密集区域偏移,压缩正常样本,使异常相对孤立,再进行高斯密度估计与马氏距离评分。该精炼不增加参数或结构修改,可叠加于任何预训练编码器。在涵盖五种成像模态(X-ray、MRI、眼底、皮肤镜、组织病理)的七个数据集上评估,该框架在四项指标中取得最高AUC,五项指标中最高平均精度,优于专用重建与扩散方法,且所有模态仅用一组固定超参数配置。结果表明,通过事后几何精炼已有表示即可实现显著提升,无需定制编码器,为实际多模态临床工作流提供实用、可扩展的AI筛查方案。
原文摘要 · Abstract (English)
Deploying AI-based anomaly detection across diverse clinical imaging settings remains challenging because most existing methods rely on modality-specific architectures, anatomical priors, or extensive retraining, limiting their use as general-purpose screening tools. One-class classification (OCC) offers a label-efficient alternative by training exclusively on normal data, but conventional two-stage pipelines fit a density estimator directly on raw pretrained embeddings, leaving substantial discriminative structure in the latent space unexploited. We introduce a training-free, modality-agnostic framework that inserts an explicit manifold-refinement stage between feature extraction and anomaly scoring. Empirical density weights, estimated via a UMAP-derived neighborhood graph, guide an iterative shift of embeddings toward locally dense regions, compacting normal samples, leaving anomalies relatively isolated prior to Gaussian density estimation and Mahalanobis-based scoring. This refinement introduces no additional trainable parameters and no architectural modification, allowing it to be layered onto any pretrained encoder. Evaluated on the MedIAnomaly benchmark across seven datasets spanning five imaging modalities (X-ray, MRI, fundus, dermatoscopy, histopathology), the framework achieves the best AUC on four datasets and the best Average Precision on five datasets among methods evaluated in the benchmark, outperforming specialized reconstruction and diffusion-based methods with a single fixed hyperparameter configuration across all modalities. These results demonstrate that meaningful gains can be achieved through post-hoc geometric refinement of existing representations rather than bespoke encoders, offering a practical and scalable AI screening framework for real-world, multi-modality clinical workflows where retraining and abnormal-case annotation are costly or infeasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。