用多层级令牌融合提升内镜图像的跨模态理解能力
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
- 通过多级分类令牌聚合增强视觉语言模型表征
- 在有限医疗数据下实现95%分类准确率与0.93召回率
- 适合低资源环境下的临床影像智能分析应用
我们提出一种面向耳鼻喉科内镜图像分析的统一视觉-语言框架,同时解决图像分类、图像到图像检索和文本到图像检索三类临床任务。不同于传统基于CNN的方法在跨模态语义捕捉上的不足,本方法采用CLIP ViT-B/16作为骨干网络,结合低秩适配、多层级分类令牌聚合和球面特征插值,实现在小样本医学数据下的高效微调,提升表征多样性与模态间语义对齐。为连接视觉输入与文本诊断语境,引入类别相关的自然语言提示,通过联合监督分类与对比学习目标引导图像编码器。在ACM MM'25 ENTRep Grand Challenge中,分类准确率与F1-score达95%,图像到图像检索的Recall@1为0.93,文本到图像检索为0.92,对应MRR分别为0.97和0.96。消融实验验证了各组件的渐进增益,证明该设计在低资源临床场景下具备鲁棒的多模态理解能力。
原文摘要 · Abstract (English)
We present a unified vision-language framework tailored for ENT endoscopy image analysis that simultaneously tackles three clinically-relevant tasks: image classification, image-to-image retrieval, and text-to-image retrieval. Unlike conventional CNN-based pipelines that struggle to capture cross-modal semantics, our approach leverages the CLIP ViT-B/16 backbone and enhances it through Low-Rank Adaptation, multi-level CLS token aggregation, and spherical feature interpolation. These components collectively enable efficient fine-tuning on limited medical data while improving representation diversity and semantic alignment across modalities. To bridge the gap between visual inputs and textual diagnostic context, we introduce class-specific natural language prompts that guide the image encoder through a joint training objective combining supervised classification with contrastive learning. We validated our framework through participation in the ACM MM'25 ENTRep Grand Challenge, achieving 95% accuracy and F1-score in classification, Recall@1 of 0.93 and 0.92 for image-to-image and text-to-image retrieval respectively, and MRR scores of 0.97 and 0.96. Ablation studies demonstrated the incremental benefits of each architectural component, validating the effectiveness of our design for robust multimodal medical understanding in low-resource clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。