arXiv:2603.10519cs.CV2026-03

通过视觉引导解耦医学图像生成,实现结构级精准控制。

Visually-Guided Controllable Medical Image Generation via Fine-Grained Semantic Disentanglement

  • 用视觉先验解耦文本语义,分离解剖结构与成像风格
  • 在三个数据集上生成质量优于现有方法,分类任务性能提升显著
  • 适合需要精细控制的医学图像生成研究者使用

医学图像合成对缓解数据稀缺和隐私问题至关重要。然而,微调通用文本到图像(T2I)模型仍具挑战性,主要源于复杂视觉细节与抽象临床文本之间的显著模态差距。此外,语义纠缠问题持续存在,粗粒度文本嵌入模糊了解剖结构与成像风格的边界,削弱了生成过程中的可控性。为此,我们提出一种视觉引导的文本解耦框架。引入跨模态潜在对齐机制,利用视觉先验将非结构化文本显式解耦为独立语义表征。随后,混合特征融合模块(HFFM)通过独立通道将这些特征注入扩散变换器(DiT),实现细粒度结构控制。在三个数据集上的实验结果表明,该方法在生成质量上超越现有方法,并显著提升下游分类任务表现。源代码已公开于 https://github.com/hx111/VG-MedGen。

原文摘要 · Abstract (English)

Medical image synthesis is crucial for alleviating data scarcity and privacy constraints. However, fine-tuning general text-to-image (T2I) models remains challenging, mainly due to the significant modality gap between complex visual details and abstract clinical text. In addition, semantic entanglement persists, where coarse-grained text embeddings blur the boundary between anatomical structures and imaging styles, thus weakening controllability during generation. To address this, we propose a Visually-Guided Text Disentanglement framework. We introduce a cross-modal latent alignment mechanism that leverages visual priors to explicitly disentangle unstructured text into independent semantic representations. Subsequently, a Hybrid Feature Fusion Module (HFFM) injects these features into a Diffusion Transformer (DiT) through separated channels, enabling fine-grained structural control. Experimental results in three datasets demonstrate that our method outperforms existing approaches in terms of generation quality and significantly improves performance on downstream classification tasks. The source code is available at https://github.com/hx111/VG-MedGen.

医学图像生成扩散模型语义解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。