用扩散模型生成真实感喉部病变图像,解决医疗数据稀缺问题。
Clinically-guided Data Synthesis for Laryngeal Lesion Detection
- 用控制网引导扩散模型生成带标注的喉镜图像。
- 仅加10%合成数据,检测率提升9%(内部)和22.1%(外部)。
- 专家难分辨真假图像,适合临床级合成数据应用。
尽管计算机辅助诊断(CADx)和检测(CADe)在多个医学领域取得进展,但在耳鼻喉科等专业领域仍受限。当前评估高度依赖医生经验,病变异质性强,活检仍是金标准但成本高、风险大。专用内镜CADx/e系统的关键瓶颈在于缺乏足够多样且标注良好的数据集。本研究提出一种新方法:利用潜空间扩散模型(LDM)结合控制网适配器,基于临床观察生成喉镜图像与标注对。该方法通过条件扩散过程生成真实、高质量、具临床意义的图像特征,涵盖多种解剖状态。实验表明,在下游检测任务中,仅增加10%合成数据,模型在内部测试中检测率提升9%,在跨域外部数据上提升22.1%。此外,5位不同经验水平的耳鼻喉科专家评估生成图像的真实性,其区分真假的能力有限,验证了图像的临床可信度。该工作为自动化喉部疾病诊断工具开发提供数据解决方案,推动合成数据在真实场景中的应用。
原文摘要 · Abstract (English)
Although computer-aided diagnosis (CADx) and detection (CADe) systems have made significant progress in various medical domains, their application is still limited in specialized fields such as otorhinolaryngology. In the latter, current assessment methods heavily depend on operator expertise, and the high heterogeneity of lesions complicates diagnosis, with biopsy persisting as the gold standard despite its substantial costs and risks. A critical bottleneck for specialized endoscopic CADx/e systems is the lack of well-annotated datasets with sufficient variability for real-world generalization. This study introduces a novel approach that exploits a Latent Diffusion Model (LDM) coupled with a ControlNet adapter to generate laryngeal endoscopic image-annotation pairs, guided by clinical observations. The method addresses data scarcity by conditioning the diffusion process to produce realistic, high-quality, and clinically relevant image features that capture diverse anatomical conditions. The proposed approach can be leveraged to expand training datasets for CADx/e models, empowering the assessment process in laryngology. Indeed, during a downstream task of detection, the addition of only 10% synthetic data improved the detection rate of laryngeal lesions by 9% when the model was internally tested and 22.1% on out-of-domain external data. Additionally, the realism of the generated images was evaluated by asking 5 expert otorhinolaryngologists with varying expertise to rate their confidence in distinguishing synthetic from real images. This work has the potential to accelerate the development of automated tools for laryngeal disease diagnosis, offering a solution to data scarcity and demonstrating the applicability of synthetic data in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。