arXiv:2504.09441cs.CVeess.IV2025-04被引 1

用动态频率平衡与医学知识引导,解决医学图像生成中的结构失真问题。

Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance

  • 通过小波变换分离高低频特征,动态调节频率以增强结构信息。
  • 在多个数据集上显著提升生成图像的结构准确性和视觉质量。
  • 适合医学影像生成、临床辅助诊断等需要高结构保真的场景。

多模态医学影像在精准全面的临床诊断中至关重要。扩散模型是生成所需医学图像的强大工具,但现有方法仍因高频信息过拟合和低频信息弱化导致解剖结构失真。为此,我们提出一种基于动态频率平衡与知识引导的新方法。首先,利用小波变换分解模型关键特征的高低频成分;随后设计动态频率平衡模块,自适应调整频率,强化全局低频结构特征,保留有效高频细节并抑制高频噪声。为应对不同医学模态间的巨大差异,构建知识引导机制,融合视觉语言模型提供的先验临床知识与视觉特征,促进解剖结构的准确生成。在多个数据集上的实验评估表明,该方法在定性与定量指标上均取得显著提升,验证了其有效性与优越性。

原文摘要 · Abstract (English)

Multimodal medical images play a crucial role in the precise and comprehensive clinical diagnosis. Diffusion model is a powerful strategy to synthesize the required medical images. However, existing approaches still suffer from the problem of anatomical structure distortion due to the overfitting of high-frequency information and the weakening of low-frequency information. Thus, we propose a novel method based on dynamic frequency balance and knowledge guidance. Specifically, we first extract the low-frequency and high-frequency components by decomposing the critical features of the model using wavelet transform. Then, a dynamic frequency balance module is designed to adaptively adjust frequency for enhancing global low-frequency features and effective high-frequency details as well as suppressing high-frequency noise. To further overcome the challenges posed by the large differences between different medical modalities, we construct a knowledge-guided mechanism that fuses the prior clinical knowledge from a visual language model with visual features, to facilitate the generation of accurate anatomical structures. Experimental evaluations on multiple datasets show the proposed method achieves significant improvements in qualitative and quantitative assessments, verifying its effectiveness and superiority.

医学图像扩散模型结构保真知识引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。