arXiv:2505.24421eess.IVcs.CV2025-05

通过多编码器增强感知学习,提升医学影像转换在不同设备下的鲁棒性。

pyMEAL: A Multi-Encoder Augmentation-Aware-Learning Toolbox for Robust Medical Image Translation

  • 用多个编码路径分别处理不同增强数据,捕捉多样化特征。
  • 在未见过的数据集上,PSNR和SSIM均优于现有方法。
  • 适合临床场景中影像设备差异大时的图像转换任务。

医学影像在临床诊断中至关重要,但人工智能驱动的影像技术仍受患者个体差异、图像伪影及成像条件变化影响,尤其在3D影像转换中因训练数据有限、扫描仪差异、成像协议和患者运动等因素导致性能受限。传统数据增强通常采用单一变换流程,忽略增强特异性,限制表征学习。为此,我们提出多编码器增强感知学习(MEAL),通过专用编码路径处理多种增强变体。研究了三种特征融合策略:编码器拼接(MEAL-CC)、融合层(MEAL-FL)和自适应控制器模块(MEAL-BD)。MEAL-BD通过动态加权增强特异性特征,在解码前保留互补表示,显著提升对临床相关变异的鲁棒性。我们在CT转T1加权MRI任务上评估,该任务在无法获取、不适宜或延迟进行MRI时具有重要临床价值。在预定义和未见测试集上,无论在几何扰动还是标准成像条件下,MEAL-BD均持续优于对比方法,取得更高峰值信噪比(PSNR)与结构相似性指数测量(SSIM)。MEAL强调结构保真度而非感知真实感,支持临床解读与下游分析,而非替代诊断级MRI,证明增强感知表征学习可提升医学影像转换的鲁棒性与临床适用性。

原文摘要 · Abstract (English)

Medical imaging plays a vital role in clinical diagnosis, yet AI-driven imaging methods remain challenged by patient variability, image artifacts, and limited robustness across acquisition conditions. Although deep learning has advanced medical image analysis, 3D image translation remains hindered by limited training data and variability arising from scanner differences, imaging protocols, and patient motion. Conventional data augmentation typically relies on a single transformation pipeline, overlooking augmentation-specific characteristics and limiting representation learning. To address these challenges, we propose Multi-Encoder Augmentation-Aware Learning (MEAL), which processes multiple augmentation variants through dedicated encoder pathways. Three feature integration strategies are investigated: encoder concatenation (MEAL-CC), fusion layer (MEAL-FL), and an adaptive controller block (MEAL-BD). By dynamically weighting augmentation-specific features before decoding, MEAL-BD preserves complementary representations and improves robustness to clinically relevant variability. We evaluate MEAL using CT-to-T1-weighted MRI translation, a clinically relevant task when MRI is unavailable, contraindicated, or delayed. Across predefined and unseen test datasets, MEAL-BD consistently outperformed competing approaches under both geometric perturbations and standard imaging conditions, achieving higher peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). By prioritizing structural fidelity over perceptual realism, MEAL supports clinical interpretation and downstream image analysis rather than replacing diagnostic MRI, demonstrating that augmentation-aware representation learning improves the robustness and clinical applicability of medical image translation.

医学影像图像转换增强学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。