arXiv:2608.23853cs.CV2026-08

LUX让内镜图像描述能精准定位病灶,提升可解释性。

LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning

论文配图:LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning
图 1 · 摘自论文原文
  • 基于病灶激活图构建病变关系图,指导生成关注具体病灶
  • 在多个指标上超越现有模型,尤其在CIDEr上提升显著
  • 适合需要可解释医疗视觉描述的临床场景

溃疡性结肠炎的内镜图像解读复杂且主观,存在评估差异和细微黏膜炎症。尽管深度学习推动了自动化分析,但多数视觉语言模型依赖全局视觉特征,忽略了病理证据的局部性和关联性,限制了临床可靠性与可解释性。我们提出LUX(Lesion-aware Unified eXplainable captioning),一种基于图结构的可解释内镜图像描述架构。LUX利用Grad-CAM和CBAM激活图构建以病灶为中心的场景图,将病理区域作为节点,并编码其空间与临床关系。这些图嵌入被融入T5解码器的交叉注意力层,使生成词汇能关注特定病灶节点而非仅全局图像特征。这实现了语言内容与病理证据的直接对齐,支持词级可解释性与关系推理。LUX在BLEU、METEOR、ROUGE-L和CIDEr等指标上优于强基线与先进模型,尤其在CIDEr上表现突出;同时减少幻觉性临床发现,通过更强的生成词与局部病理区域对应,提升了病灶级定位能力。

原文摘要 · Abstract (English)

The interpretation of endoscopic imagery in ulcerative colitis is complex and subjective, with variability in human assessment and subtle mucosal inflammation. Although deep learning has advanced automated analysis, most vision-language models rely on global visual embeddings that overlook the localized and relational nature of pathological evidence, limiting clinical reliability and interpretability. We introduce LUX (Lesion-aware Unified eXplainable captioning), a graph-conditioned vision-language architecture for explainable endoscopic image captioning. LUX constructs a lesion-centric scene graph from Grad-CAM and CBAM activation maps, representing pathological regions as nodes and encoding their spatial and clinical relationships. These graph embeddings are integrated into the cross-attention layers of a T5 decoder, enabling generated words to attend to specific lesion nodes rather than only to global image features. This provides direct alignment between linguistic content and pathological evidence, supporting token-level interpretability and relational reasoning. LUX outperforms strong baseline and state-of-the-art medical captioning models across BLEU, METEOR, ROUGE-L, and CIDEr, with particularly strong gains in CIDEr. It also reduces hallucinated clinical findings and improves lesion-level grounding through stronger correspondence between generated tokens and localized pathological regions.

可解释性医学图像描述图神经网络内镜分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。