arXiv:2412.17251cs.CVcs.LG2024-12中稿 · presentation at th…被引 10

用引导式注意力提升眼底图像描述生成质量

GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning

  • 引入引导式上下文自注意力融合多模态信息
  • 在有限标注数据下提升BLEU@4达0.023
  • 适合医学影像报告生成与临床辅助诊断

眼底图像分析对眼病诊断至关重要,但受限于图像质量差异和病理多样性,尤其在标注数据稀缺的情况下,生成准确医疗报告仍具挑战。以往基于Transformer的模型难以在弱监督下有效融合视觉与文本信息。为此,我们提出一种新型多模态视觉语言变压器(GCS-M3VLT),通过引导式上下文自注意力机制整合视觉与文本特征,即使在数据稀缺场景下也能捕捉细粒度细节与全局临床语境。在DeepEyeNet数据集上的大量实验表明,该模型实现0.023的BLEU@4提升,并带来显著的定性改进,验证了其生成全面医疗描述的有效性。

原文摘要 · Abstract (English)

Retinal image analysis is crucial for diagnosing and treating eye diseases, yet generating accurate medical reports from images remains challenging due to variability in image quality and pathology, especially with limited labeled data. Previous Transformer-based models struggled to integrate visual and textual information under limited supervision. In response, we propose a novel vision-language model for retinal image captioning that combines visual and textual features through a guided context self-attention mechanism. This approach captures both intricate details and the global clinical context, even in data-scarce scenarios. Extensive experiments on the DeepEyeNet dataset demonstrate a 0.023 BLEU@4 improvement, along with significant qualitative advancements, highlighting the effectiveness of our model in generating comprehensive medical captions.

眼底图像多模态自注意力医学报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。