通过因果解耦生成更真实多样的罕见病医学影像。
Causal Disentanglement for Robust Long-tail Medical Image Generation
- 用因果解耦分离病理与结构特征,确保独立可控。
- 结合文本引导扩散模型,生成多样且符合报告的病变图像。
- 优化初始噪声提升长尾类别生成质量,适合医疗数据生成场景。
反事实医学图像生成能缓解数据稀缺并提升图像可解释性。但医学图像病理特征复杂、数据分布不均,从有限数据中生成高质量、多样化的图像极具挑战。为充分挖掘有限数据中的解剖结构信息,生成结构更稳定的图像,避免失真或不一致,本文提出一种新型医学图像生成框架。该框架基于因果解耦,分离独立的病理与结构特征,并利用文本引导建模病理特征以调控反事实图像生成。首先,通过因果解耦实现特征分离,并引入组监督确保病理与身份特征独立。其次,采用由病灶描述引导的扩散模型建模病理特征,实现多样化反事实图像生成;同时,借助大语言模型从医学报告中提取病灶严重程度和位置信息以提高准确性。此外,通过初始噪声优化提升潜在扩散模型在长尾类别上的表现。
原文摘要 · Abstract (English)
Counterfactual medical image generation effectively addresses data scarcity and enhances the interpretability of medical images. However, due to the complex and diverse pathological features of medical images and the imbalanced class distribution in medical data, generating high-quality and diverse medical images from limited data is significantly challenging. Additionally, to fully leverage the information in limited data, such as anatomical structure information and generate more structurally stable medical images while avoiding distortion or inconsistency. In this paper, in order to enhance the clinical relevance of generated data and improve the interpretability of the model, we propose a novel medical image generation framework, which generates independent pathological and structural features based on causal disentanglement and utilizes text-guided modeling of pathological features to regulate the generation of counterfactual images. First, we achieve feature separation through causal disentanglement and analyze the interactions between features. Here, we introduce group supervision to ensure the independence of pathological and identity features. Second, we leverage a diffusion model guided by pathological findings to model pathological features, enabling the generation of diverse counterfactual images. Meanwhile, we enhance accuracy by leveraging a large language model to extract lesion severity and location from medical reports. Additionally, we improve the performance of the latent diffusion model on long-tailed categories through initial noise optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。