arXiv:2603.26393eess.IVcs.CV2026-03

用轻量适配让冻结的单模模型搞定多模态医学图像配准。

Adapting Frozen Mono-modal Backbones for Multi-modal Registration via Contrast-Agnostic Instance Optimization

  • 通过无对比的表征生成与实例优化,实现测试时动态适配。
  • 在多模态和域外数据上均超越预训练单模基线,排名前四。
  • 无需全网微调,适合资源受限场景下的鲁棒配准任务。

可变形图像配准在医学图像分析中仍是核心挑战,尤其在模态间强度分布差异显著的情况下。深度学习方法虽能高效预测,但在测试时面对分布偏移常表现不佳。直接全网微调虽有效,但对Transformer或深层U-Net等现代架构在3D场景下内存与运行时间开销巨大。此外,简单微调在剧烈域偏移下易导致性能下降。本文提出一种融合冻结预训练单模配准模型与轻量级适配管道的配准框架,利用基于无对比表征生成与精炼模块的风格迁移,在测试时通过实例优化弥合模态与域间差距。该设计与单模骨干选择无关,避免全微调开销,同时具备适应未见域的灵活性。在Learn2Reg 2025 LUMIR验证集上评估,该方法在多模态子集上排名第二,域外子集第三,总体Dice分数第四。结果表明,结合冻结单模模型与模态适配及轻量实例优化,是实现稳健多模态配准的有效路径。

原文摘要 · Abstract (English)

Deformable image registration remains a central challenge in medical image analysis, particularly under multi-modal scenarios where intensity distributions vary significantly across scans. While deep learning methods provide efficient feed-forward predictions, they often fail to generalize robustly under distribution shifts at test time. A straightforward remedy is full network fine-tuning, yet for modern architectures such as Transformers or deep U-Nets, this adaptation is prohibitively expensive in both memory and runtime when operating in 3D. Meanwhile, the naive fine-tuning struggles more with potential degradation in performance in the existence of drastic domain shifts. In this work, we propose a registration framework that integrates a frozen pretrained \textbf{mono-modal} registration model with a lightweight adaptation pipeline for \textbf{multi-modal} image registration. Specifically, we employ style transfer based on contrast-agnostic representation generation and refinement modules to bridge modality and domain gaps with instance optimization at test time. This design is orthogonal to the choice of backbone mono-modal model, thus avoids the computational burden of full fine-tuning while retaining the flexibility to adapt to unseen domains. We evaluate our approach on the Learn2Reg 2025 LUMIR validation set and observe consistent improvements over the pretrained state-of-the-art mono-modal backbone. In particular, the method ranks second on the multi-modal subset, third on the out-of-domain subset, and achieves fourth place overall in Dice score. These results demonstrate that combining frozen mono-modal models with modality adaptation and lightweight instance optimization offers an effective and practical pathway toward robust multi-modal registration.

图像配准多模态轻量适配医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。