arXiv:2601.09209cs.CV2026-01AAAI被引 1

无需配对图像,用分组知识蒸馏提升胃肠道病变分类准确率

Pairing-free Group-level Knowledge Distillation for Robust Gastrointestinal Lesion Classification in White-Light Endoscopy

  • 基于分组级知识蒸馏,不依赖成对的白光与窄带图像
  • 在4个临床数据集上提升AUC,最高达3.3%相对改进
  • 适合缺乏配对数据的医疗影像研究者使用

白光成像(WLI)是内镜癌症筛查的标准,而窄带成像(NBI)提供更优诊断细节。现有方法需依赖同一病灶的成对WLI-NBI图像,成本高且难以实现,导致大量临床数据无法利用。本文提出无配对分组级知识蒸馏(PaGKD)框架,首次实现仅用非配对的WLI和NBI数据进行跨模态学习。其核心为两个互补模块:(1) 分组原型蒸馏(GKD-Pro)通过共享病变感知查询提取模态不变语义原型;(2) 分组密集蒸馏(GKD-Den)利用激活导出的关系图引导组感知注意力,实现密集跨模态对齐。二者共同保证全局语义一致性和局部结构连贯性,无需图像级对应关系。在四个临床数据集上的实验表明,PaGKD显著优于现有方法,相对AUC提升分别为3.3%、1.1%、2.8%和3.2%,为非配对跨模态学习开辟新路径。

原文摘要 · Abstract (English)

White-Light Imaging (WLI) is the standard for endoscopic cancer screening, but Narrow-Band Imaging (NBI) offers superior diagnostic details. A key challenge is transferring knowledge from NBI to enhance WLI-only models, yet existing methods are critically hampered by their reliance on paired NBI-WLI images of the same lesion, a costly and often impractical requirement that leaves vast amounts of clinical data untapped. In this paper, we break this paradigm by introducing PaGKD, a novel Pairing-free Group-level Knowledge Distillation framework that that enables effective cross-modal learning using unpaired WLI and NBI data. Instead of forcing alignment between individual, often semantically mismatched image instances, PaGKD operates at the group level to distill more complete and compatible knowledge across modalities. Central to PaGKD are two complementary modules: (1) Group-level Prototype Distillation (GKD-Pro) distills compact group representations by extracting modality-invariant semantic prototypes via shared lesion-aware queries; (2) Group-level Dense Distillation (GKD-Den) performs dense cross-modal alignment by guiding group-aware attention with activation-derived relation maps. Together, these modules enforce global semantic consistency and local structural coherence without requiring image-level correspondence. Extensive experiments on four clinical datasets demonstrate that PaGKD consistently and significantly outperforms state-of-the-art methods, achieving relative AUC improvements of 3.3%, 1.1%, 2.8%, and 3.2%, respectively, establishing a new direction for cross-modal learning from unpaired data.

医学影像知识蒸馏跨模态学习无配对数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。