arXiv:2607.08397cs.CV2026-07

解决内镜图像中开放词汇的细粒度语义分割难题

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

论文配图:Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation
图 1 · 摘自论文原文
  • 通过属性检索实现开放词汇的内镜图像语义定位
  • 在真实与模拟数据上均达当前最优性能
  • 专为内镜场景设计,适合医疗视觉研究者

指代图像分割(RIS)旨在根据自然语言描述分割图像区域,实现细粒度、可控制的视觉理解。将RIS拓展至内镜影像时面临独特挑战:高质量标注稀缺,且图像-文本关系复杂且领域特定。尽管近期视觉-语言模型展现出强跨域对齐能力,但在内镜场景下仍难以捕捉细粒度文本线索,导致性能不佳且泛化能力有限。为此,我们构建了首个内镜领域的大规模基准数据集ReferEndoscopy。基于该数据集,提出基于属性检索的开放词汇内镜组合指代分割框架AR-ERIS。AR-ERIS在精心构建的ReferEndoscopy数据集上预训练,实现了在模拟和真实内镜数据上的最佳性能,具备优异泛化能力。数据集与代码将在评审完成后公开。

原文摘要 · Abstract (English)

Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique challenges, including scarce high-quality annotations and complex, domain-specific image-text relationships. Although recent vision-language models demonstrate strong cross-domain alignment, they often fail to capture fine-grained textual cues in endoscopic settings, resulting in suboptimal performance and limited generalization. To address these challenges, we introduce ReferEndoscopy, a large-scale benchmark for RIS in the endoscopy field. Building on this dataset, we propose the Attribute Retrieval-based Endoscopic-RIS (AR-ERIS) framework for open-vocabulary endoscopic compositional referring segmentation. AR-ERIS leverages attribute retrieval for open-vocabulary endoscopic compositional referring segmentation and is pretrained on the curated ReferEndoscopy dataset, achieving state-of-the-art performance with strong generalization across both simulated and real-world endoscopic data. The dataset and code will be publicly released upon completion of the review process.

图像分割内镜视觉开放词汇多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。