用文本控制生成多样肠镜图像,提升息肉分类准确率
Diverse Image Generation with Diffusion Models and Cross Class Label Learning for Polyp Classification
- 引入跨类别标签学习,减少对标注数据的依赖
- 生成涵盖不同病理类型、成像模式的合成图像,提升分类性能
- 适合医学影像生成与少样本分类研究者使用
病理诊断是决定结直肠癌(CRC)最佳治疗方案的关键环节。结肠息肉作为CRC前兆,可分为腺瘤性和增生性两大类。为实现精准分类和早期诊断,结肠镜检查结合窄带成像与白光成像等技术广泛应用。然而现有分类方法多依赖单一成像模态,受限于数据稀缺,性能有限。近年来,生成式人工智能在克服此类问题方面展现潜力。尽管文本提示与图像控制机制已用于生成高质量图像,但在结肠镜领域,尤其是文本提示控制尚未被探索。此外,昂贵的类别标签获取限制了相关研究。为此,我们提出新模型PathoPolyp-Diff,可生成受文本控制、具有病理多样性、成像模态多样性和图像质量差异的合成图像。通过跨类别标签学习,模型能从其他类别中提取特征,减轻标注负担。实验结果表明,在公开数据集上,平衡准确率提升达7.91%;视频级分析中,跨类别标签学习实现高达18.33%的统计显著提升。代码已开源。
原文摘要 · Abstract (English)
Pathologic diagnosis is a critical phase in deciding the optimal treatment procedure for dealing with colorectal cancer (CRC). Colonic polyps, precursors to CRC, can pathologically be classified into two major types: adenomatous and hyperplastic. For precise classification and early diagnosis of such polyps, the medical procedure of colonoscopy has been widely adopted paired with various imaging techniques, including narrow band imaging and white light imaging. However, the existing classification techniques mainly rely on a single imaging modality and show limited performance due to data scarcity. Recently, generative artificial intelligence has been gaining prominence in overcoming such issues. Additionally, various generation-controlling mechanisms using text prompts and images have been introduced to obtain visually appealing and desired outcomes. However, such mechanisms require class labels to make the model respond efficiently to the provided control input. In the colonoscopy domain, such controlling mechanisms are rarely explored; specifically, the text prompt is a completely uninvestigated area. Moreover, the unavailability of expensive class-wise labels for diverse sets of images limits such explorations. Therefore, we develop a novel model, PathoPolyp-Diff, that generates text-controlled synthetic images with diverse characteristics in terms of pathology, imaging modalities, and quality. We introduce cross-class label learning to make the model learn features from other classes, reducing the burdensome task of data annotation. The experimental results report an improvement of up to 7.91% in balanced accuracy using a publicly available dataset. Moreover, cross-class label learning achieves a statistically significant improvement of up to 18.33% in balanced accuracy during video-level analysis. The code is available at https://github.com/Vanshali/PathoPolyp-Diff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。