让医学影像融合更懂任务:自动优化肿瘤边界,提升分割精度。
Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization
- 用下游分割任务反向指导图像融合,实现任务驱动优化。
- 在多种模态下显著超越现有双通道分割方法,保持高精度边缘细节。
- 融合结果兼具物理真实性和可解释性,适合临床部署与医生信任。
多模态医学图像融合传统上以人眼视觉感知为目标,追求通用对比度和结构保真度。然而,当这些视觉上优美的融合图像用于自动化临床流程时,这种视觉-语义差异会导致任务无关的特征退化,无意中模糊关键的高频肿瘤边界。为弥合这一语义鸿沟,我们提出Fuse4Seg,将多模态融合重新建模为医疗分割任务下的协同双层优化问题。融合主模型不再依赖固定视觉指标,而是根据下游分割辅模型回传的语义梯度动态调整特征提取策略。为确保物理保真度与语义实用性并重,设计了频域解耦架构,并通过频域分解损失(Frequency Decomposition Loss)和空间梯度损失(Spatial Gradient Loss)严格正则化。该显式物理锚点有效防止解剖畸变,保障任务关键细节无损保留。大量实验表明,我们的任务感知单通道融合先验能无缝泛化于多样多尺度模态。更令人振奋的是,其分割性能显著超越当前双通道分割最先进方法,同时明确提供可读性强的“透明盒”物理图像,增强临床可视化解释力与信任度。
原文摘要 · Abstract (English)
Multi-modal medical image fusion is traditionally optimized for human visual perception, aiming to maximize generic contrast and structural fidelity. However, when these visually pleasing fused images are deployed in automated clinical workflows, this visual-semantic discrepancy causes task-agnostic feature degradation, inadvertently smoothing out critical, high-frequency tumor boundaries. To bridge this semantic gap, we propose Fuse4Seg, a novel framework that reformulates multi-modal fusion as a cooperative bi-level optimization problem with medical segmentation. Rather than relying on rigid visual metrics, our fusion leader dynamically updates its feature extraction strategy driven directly by semantic gradients backpropagated from the downstream segmentation follower. To guarantee robust physical fidelity alongside semantic utility, we design a frequency-decoupled architecture stringently regularized by a Frequency Decomposition Loss and a Spatial Gradient Loss. This explicit physical anchor prevents anatomical distortion and ensures the lossless preservation of task-critical details. Extensive experiments demonstrate that our task-aware, single-channel fused prior generalizes seamlessly across diverse multi-scale modalities. More impressively, it remarkably surpasses contemporary dual-channel segmentation state-of-the-arts while explicitly providing a readable, "glass-box" physical image to foster clinical visual interpretability and trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。