构建智能肠镜三大基础工具,推动多模态医学应用
Frontiers in Intelligent Colonoscopy
- 构建大规模多模态指令数据集ColonINST与专用模型ColonGPT
- 覆盖分类、检测、分割、视觉语言理解四类任务,提升场景感知能力
- 开源平台实时更新,助力临床研究与技术迭代
肠镜检查是目前最灵敏的结直肠癌筛查手段。本文系统探讨智能肠镜技术前沿及其在多模态医疗中的潜在应用。通过四项任务——分类、检测、分割和视觉语言理解——评估当前以数据和模型为中心的研究格局,揭示领域内仍存在显著挑战,尤其在多模态研究方面尚有广阔探索空间。为迎接多模态时代,本文提出三项基础建设:大规模多模态指令微调数据集ColonINST、专为肠镜设计的多模态语言模型ColonGPT,以及一个综合性多模态评测基准。为持续追踪该快速发展的领域,项目提供公开网站(https://github.com/ai4colonoscopy/IntelliScope)进行最新进展更新。
原文摘要 · Abstract (English)
Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal medical applications. With this goal, we begin by assessing the current data-centric and model-centric landscapes through four tasks for colonoscopic scene perception, including classification, detection, segmentation, and vision-language understanding. This assessment enables us to identify domain-specific challenges and reveals that multimodal research in colonoscopy remains open for further exploration. To embrace the coming multimodal era, we establish three foundational initiatives: a large-scale multimodal instruction tuning dataset ColonINST, a colonoscopy-designed multimodal language model ColonGPT, and a multimodal benchmark. To facilitate ongoing monitoring of this rapidly evolving field, we provide a public website for the latest updates: https://github.com/ai4colonoscopy/IntelliScope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。