arXiv:2508.07028cs.CV2025-08被引 1

用图神经网络和大模型评估,提升肠镜图像中息肉分割精度。

Large Language Model Evaluated Stand-alone Attention-Assisted Graph Neural Network with Spatial and Structural Information Interaction for Precise Endoscopic Image Segmentation

  • 融合空间与结构图信息,捕捉息肉的拓扑与连接特征。
  • 在五个指标上达到当前最佳,边界保持更精准。
  • 首次用大模型自动评估分割质量,适合医疗AI研究者。

准确分割肠镜图像中的息肉对早期结直肠癌检测至关重要。但因与周围黏膜对比度低、反光强、边界模糊,该任务仍具挑战。为此,本文提出FOCUS-Med(基于注意力感知的时空图融合息肉分割),通过双图卷积网络(Dual-GCN)模块捕获上下文空间与拓扑结构依赖关系,增强模型对复杂形状与细微边界的分辨能力。引入独立位置融合的自注意力机制以强化全局上下文建模,并采用可训练加权快速归一化融合策略,实现编码器与解码器间多尺度信息高效聚合。特别地,首次利用大语言模型(LLM)对分割结果进行定性评估。在多个公开数据集上的实验表明,FOCUS-Med在五项关键指标上均达到当前最优性能,验证了其在人工智能辅助肠镜中的有效性与临床潜力。

原文摘要 · Abstract (English)

Accurate endoscopic image segmentation on the polyps is critical for early colorectal cancer detection. However, this task remains challenging due to low contrast with surrounding mucosa, specular highlights, and indistinct boundaries. To address these challenges, we propose FOCUS-Med, which stands for Fusion of spatial and structural graph with attentional context-aware polyp segmentation in endoscopic medical imaging. FOCUS-Med integrates a Dual Graph Convolutional Network (Dual-GCN) module to capture contextual spatial and topological structural dependencies. This graph-based representation enables the model to better distinguish polyps from background tissues by leveraging topological cues and spatial connectivity, which are often obscured in raw image intensities. It enhances the model's ability to preserve boundaries and delineate complex shapes typical of polyps. In addition, a location-fused stand-alone self-attention is employed to strengthen global context integration. To bridge the semantic gap between encoder-decoder layers, we incorporate a trainable weighted fast normalized fusion strategy for efficient multi-scale aggregation. Notably, we are the first to introduce the use of a Large Language Model (LLM) to provide detailed qualitative evaluations of segmentation quality. Extensive experiments on public benchmarks demonstrate that FOCUS-Med achieves state-of-the-art performance across five key metrics, underscoring its effectiveness and clinical potential for AI-assisted colonoscopy.

医学图像图神经网络大模型评估息肉分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。