arXiv:2512.05571cs.CV2025-12被引 1

用扩散模型提取三维医学影像特征,无需训练即可精准匹配解剖结构。

MedDIFT: Multi-Scale Diffusion-Based Correspondence in 3D Medical Imaging

  • 基于预训练扩散模型的多尺度特征作为体素描述符
  • 在肺部CT数据集上实现零训练对应匹配,准确率显著提升
  • 适合需快速配准的临床场景,如病变追踪与手术导航

医学影像间的空间对应关系对纵向分析、病灶追踪和图像引导干预至关重要。传统配准方法依赖局部强度相似性,难以捕捉全局语义结构,在低对比度或解剖可变区域常出现错配。扩散模型的中间表示蕴含丰富的几何与语义信息。本文提出MedDIFT,一种无需训练的3D对应框架,利用预训练潜在医学扩散模型的多尺度特征作为体素描述符,通过余弦相似度匹配,并可选加入局部搜索先验。在公开肺部CT数据集上,MedDIFT展现出无需任务特定训练即可识别解剖对应的能力。消融实验表明,多层级特征融合与适度扩散噪声可提升性能。代码已开源。

原文摘要 · Abstract (English)

Accurate spatial correspondence between medical images is essential for longitudinal analysis, lesion tracking, and image-guided interventions. Medical image registration methods rely on local intensity-based similarity measures, which fail to capture global semantic structure and often yield mismatches in low-contrast or anatomically variable regions. Recent advances in diffusion models suggest that their intermediate representations encode rich geometric and semantic information. We present MedDIFT, a training-free 3D correspondence framework that leverages multi-scale features from a pretrained latent medical diffusion model as voxel descriptors. MedDIFT fuses diffusion activations into rich voxel-wise descriptors and matches them via cosine similarity, with an optional local-search prior. On a publicly available lung CT dataset, MedDIFT shows promising capability in identifying anatomical correspondence without requiring any task-specific model training. Ablation experiments confirm that multi-level feature fusion and modest diffusion noise improve performance. Code is available online.

医学影像扩散模型图像配准无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。