用医生看片视线和影像组学特征,让扩散模型生成更真实的医学图像。
RadGazeGen: Radiomics and Gaze-guided Medical Image Generation using Diffusion Models
- 结合放射科医生注视轨迹与影像组学特征作为控制信号
- 在REFLACX数据集上生成图像质量高且多样性好
- 适合医学图像合成、辅助诊断研究者使用
本文提出RadGazeGen框架,通过整合放射科医生的注视模式和影像组学特征图作为控制信号,指导文本到图像扩散模型生成高保真医学图像。尽管文本到图像扩散模型已取得进展,但文本描述常难以传达疾病特异性细节,导致生成图像临床准确性不足。解剖结构、病变纹理及位置对生成真实图像至关重要,其保真度直接影响疾病诊断与治疗评估等下游任务。医生注视模式反映了细微病变特征与空间分布,影像组学特征则提供亚视觉层面的疾病表型信息。本研究将两者联合作为控制输入,以生成解剖准确且病灶感知的医学图像。在REFLACX数据集上评估了生成图像的质量与多样性;为验证临床适用性,进一步在CheXpert测试集(n=500)上进行分类性能测试,并在MIMIC-CXR-LT测试集(n=23,550)上评估长尾学习表现。
原文摘要 · Abstract (English)
In this work, we present RadGazeGen, a novel framework for integrating experts' eye gaze patterns and radiomic feature maps as controls to text-to-image diffusion models for high fidelity medical image generation. Despite the recent success of text-to-image diffusion models, text descriptions are often found to be inadequate and fail to convey detailed disease-specific information to these models to generate clinically accurate images. The anatomy, disease texture patterns, and location of the disease are extremely important to generate realistic images; moreover the fidelity of image generation can have significant implications in downstream tasks involving disease diagnosis or treatment repose assessment. Hence, there is a growing need to carefully define the controls used in diffusion models for medical image generation. Eye gaze patterns of radiologists are important visuo-cognitive information, indicative of subtle disease patterns and spatial location. Radiomic features further provide important subvisual cues regarding disease phenotype. In this work, we propose to use these gaze patterns in combination with standard radiomics descriptors, as controls, to generate anatomically correct and disease-aware medical images. RadGazeGen is evaluated for image generation quality and diversity on the REFLACX dataset. To demonstrate clinical applicability, we also show classification performance on the generated images from the CheXpert test set (n=500) and long-tailed learning performance on the MIMIC-CXR-LT test set (n=23550).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。