arXiv:2603.14559cs.CVcs.AI2026-03被引 4

首个融合双评分标准与专家描述的结肠镜溃疡性结肠炎数据集

A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy

  • 多中心、多分辨率数据集,含专家标注的MES与UCEIS评分
  • 首次实现基于图像的临床描述自动生成,支持双评分任务
  • 适合医疗影像算法研究者和临床辅助诊断系统开发者

溃疡性结肠炎(UC)是一种慢性黏膜炎症性疾病,患者患结直肠癌风险升高。结肠镜检查是评估疾病活动性的金标准,报告通常依赖标准化内镜评分指标,最常用的是 Mayo 内镜评分(MES,0–3分)和溃疡性结肠炎内镜严重程度指数(UCEIS,0–8分),分数越高表示病情越重。然而,自动化预测这些评分的计算方法仍受限于缺乏公开的专家标注数据集及稳健基准测试。同时,尽管图像字幕生成在计算机视觉中已成熟,但针对UC图像生成临床可解释描述的研究仍严重不足。不同中心间内镜设备和操作流程差异也凸显了多中心数据对算法鲁棒性和泛化能力的重要性。本文构建了一个经专家验证的多中心、多分辨率数据集,包含双重评分(MES与UCEIS)标签及详细临床描述。据我们所知,这是首个将双评分体系用于分类任务并结合专家生成字幕的综合性数据集,为开发临床有意义的多模态算法开辟新路径。此外,还提供了基于卷积神经网络、视觉变压器、混合模型及主流多模态视觉-语言字幕算法的基准测试。

原文摘要 · Abstract (English)

Ulcerative colitis (UC) is a chronic mucosal inflammatory condition that places patients at increased risk of colorectal cancer. Colonoscopic surveillance remains the gold standard for assessing disease activity, and reporting typically relies on standardised endoscopic scoring metrics. The most widely used is the Mayo Endoscopic Score (MES), with some centres also adopting the Ulcerative Colitis Endoscopic Index of Severity (UCEIS). Both are descriptive assessments of mucosal inflammation (MES: 0 to 3; UCEIS: 0 to 8), where higher values indicate more severe disease. However, computational methods for automatically predicting these scores remain limited, largely due to the lack of publicly available expert-annotated datasets and the absence of robust benchmarking. There is also a significant research gap in generating clinically meaningful descriptions of UC images, despite image captioning being a well-established computer vision task. Variability in endoscopic systems and procedural workflows across centres further highlights the need for multi-centre datasets to ensure algorithmic robustness and generalisability. In this work, we introduce a curated multi-centre, multi-resolution dataset that includes expert-validated MES and UCEIS labels, alongside detailed clinical descriptions. To our knowledge, this is the first comprehensive dataset that combines dual scoring metrics for classification tasks with expert-generated captions describing mucosal appearance and clinically accepted reasoning for image captioning. This resource opens new opportunities for developing clinically meaningful multimodal algorithms. In addition to the dataset, we also provide benchmarking using convolutional neural networks, vision transformers, hybrid models, and widely used multimodal vision-language captioning algorithms.

结肠镜多模态评分系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。