arXiv:2607.20598eess.IV2026-07

用三路联合适配,让大模型更好生成医学超分辨率图像。

MedDiT4SR: Tri-Stream Joint Adaptation of Pre-Trained Diffusion Transformers for Medical Image Super-Resolution

论文配图:MedDiT4SR: Tri-Stream Joint Adaptation of Pre-Trained Diffusion Transformers for Medical Image Super-Resolution
图 1 · 摘自论文原文
  • 将低分辨率图、噪声潜码和文本信息融合进同一扩散变压器块中
  • 引入适配器抑制插值冗余,提升细节还原能力
  • 适合医疗影像重建,尤其在跨数据集场景下表现优越

医学图像超分辨率(MedSR)需从退化图像中恢复精细解剖结构,同时避免生成模型引入不合理的细节。大规模预训练多模态扩散变换器虽具备强视觉先验,但其在MedSR任务中的适配仍具挑战性。传统ControlNet式适配中,低分辨率(LR)图像作为外部条件单向注入去噪流,导致LR解剖证据无法与演化中的去噪和语义表示协同更新。本文提出MedDiT4SR,一种三路联合适配框架,将LR图像、噪声潜码和文本表示整合至同一多模态扩散变换器块中。为补充全局令牌交互,引入超分辨率适配器(SR Adapter),聚合依赖尺度的局部令牌并抑制插值引起的冗余。进一步提出语义对齐精炼器(SA Refiner),利用提示引导的语义信息校准局部低分辨率响应。在域内及同模态跨数据集设置下的实验表明,该方法能有效将大规模预训练扩散变换器模型适配至多种成像领域的医学图像超分辨率任务。

原文摘要 · Abstract (English)

Medical image super-resolution (MedSR) requires recovering fine anatomical structures from degraded observations while avoiding unsupported details introduced by generative priors. Large-scale pre-trained multimodal diffusion transformers provide strong visual priors, but their adaptation to MedSR remains non-trivial. In conventional ControlNet-style adaptation, the low-resolution (LR) image is processed as an external condition and injected into the denoising stream through one-way connections. Consequently, LR anatomical evidence cannot be jointly updated with the evolving denoising and semantic representations. We propose MedDiT4SR, a tri-stream adaptation framework that integrates the LR, noisy latent, and text representations into the same multimodal diffusion-transformer blocks. To complement global token interaction, we introduce a Super-Resolution Adapter (SR Adapter) that aggregates scale-dependent local tokens and suppresses interpolation-induced redundancy. We further propose a Semantic Alignment Refiner (SA Refiner) that calibrates local LR responses using prompt-conditioned semantic information. Experiments under both in-domain and within-modality cross-dataset settings demonstrate the effectiveness of adapting large-scale pre-trained DiT models to medical image super-resolution across diverse imaging domains.

医学图像超分辨率扩散模型三路融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。