构建首个结肠镜多模态推理数据集,提升智能结肠镜临床决策能力。
Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
- 构建包含110万+问答的ColonVQA数据集,覆盖76种临床发现
- 提出ColonR1模型,在数据稀缺下准确率达56.61%,比微调高25.22%
- 通过多智能体辩论生成临床推理数据,推动从理解到决策跃迁
本研究提出Colon-X,一项推进结肠镜多模态智能的开放计划。首先构建了目前最全面的结肠镜多模态数据集ColonVQA,包含超过110万条视觉问答数据,覆盖76种临床发现和18项多模态任务。为探索从多模态理解到临床推理的演进,我们系统评估22个多模态大模型在人类扰动下的泛化能力,发现主流模型输出仍不可靠。为此,我们设计基于多智能体辩论的ColonReason推理数据集,并开发首个采用R1范式、通过任务自适应奖励与梯度稳定策略优化的ColonR1模型。在数据稀缺条件下,其整体准确率达到56.61%,显著优于监督微调方法(高出25.22%),建立了新的多模态结肠镜分析推理基准。所有数据与模型资源公开于https://github.com/ai4colonoscopy/Colon-X。
原文摘要 · Abstract (English)
In this study, we present Colon-X, an open initiative aimed at advancing multimodal intelligence in colonoscopy. We begin by constructing ColonVQA, the most comprehensive multimodal dataset ever built for colonoscopy, featuring over 1.1M+ visual question answering entries across 76 clinical findings and 18 multimodal tasks. Beyond serving as a community-wide data foundation, we further investigate a critical yet underexplored transition in colonoscopy - evolving from multimodal understanding to clinical reasoning: (a) To capture the current landscape of multimodal understanding behaviors, we systematically assess the generalizability of 22 multimodal large language models and examine their reliability under human-induced perturbations. The results reveal that clinical outputs from leading MLLMs remain far from robust and trustworthy. (b) To narrow this gap, we further explore reasoning-centric intelligence tailored for colonoscopy. Specifically, we curate ColonReason, a clinically grounded reasoning dataset annotated through a multi-agent debating pipeline, and develop ColonR1, the first R1-styled model that mitigates reward information collapse through task-adaptive rewards and gradient-stable policy optimization. Under data-scarce conditions, our ColonR1 achieves 56.61% overall accuracy, outperforming supervised fine-tuning by 25.22%, and sets a new reasoning-enabled baseline for multimodal colonoscopy analysis. All data and model resources are publicly available at https://github.com/ai4colonoscopy/Colon-X.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。