arXiv:2510.03543cs.CV2025-10被引 2

用AI自动生成胃肠镜报告,减轻医生负担。

From Scope to Script: An Automated Report Generation Model for Gastrointestinal Endoscopy

  • 分两阶段训练:先学图像-文本对,再微调生成报告。
  • 基于视觉语言模型,能准确提取内镜发现。
  • 适合临床医生和医学影像系统开发者使用。

胃肠道(GI)内镜检查如胃十二指肠镜(EGD)和结肠镜在诊断与管理胃肠道疾病中至关重要。然而,这些操作伴随的文档工作量极大,加剧了消化科医生的工作负担,导致临床流程低效和职业倦怠。为解决此问题,我们提出一种新型自动化报告生成模型,采用基于Transformer的视觉编码器与文本解码器,在两阶段训练框架下运行。第一阶段,双模块在图像/文本标题对上预训练,以捕捉通用视觉-语言特征;第二阶段,在图像/报告对上微调,生成具有临床意义的检查发现。该方法不仅简化了文档流程,还具有降低医生工作负荷、提升患者照护质量的潜力。

原文摘要 · Abstract (English)

Endoscopic procedures such as esophagogastroduodenoscopy (EGD) and colonoscopy play a critical role in diagnosing and managing gastrointestinal (GI) disorders. However, the documentation burden associated with these procedures place significant strain on gastroenterologists, contributing to inefficiencies in clinical workflows and physician burnout. To address this challenge, we propose a novel automated report generation model that leverages a transformer-based vision encoder and text decoder within a two-stage training framework. In the first stage, both components are pre-trained on image/text caption pairs to capture generalized vision-language features, followed by fine-tuning on images/report pairs to generate clinically meaningful findings. Our approach not only streamlines the documentation process but also holds promise for reducing physician workload and improving patient care.

医疗AI报告生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。