用眼球注视数据辅助医学影像分割,提升精度并减少人工标注时间。
Gaze-Assisted Medical Image Segmentation
- 利用医生看图时的眼动数据作为提示,微调MedSAM模型进行分割修正。
- 在120例腹部CT上实现16个器官平均Dice达90.5%,优于现有模型。
- 适合临床影像分析场景,尤其适用于需高效精准分割的放射治疗规划。
患者器官的标注是放射治疗规划等诊疗流程中的关键环节。手动标注耗时费力,而现有自动化方法尚未达到临床可用水平。本文探索以人类眼动作为交互输入,实现半监督医学图像分割修正。具体地,我们基于人体眼动数据(来自阅读腹部影像)对公开的医学图像分割模型MedSAM进行微调,使其利用眼动信息作为提示。模型在公开的WORD数据库上验证,该库包含120例腹部CT扫描及16个腹部器官标注。结果显示,结合眼动信息的MedSAM在16个器官上的平均Dice系数达90.5%,显著优于nnUNetV2(85.8%)、ResUNet(86.7%)和原始MedSAM(81.7%)。
原文摘要 · Abstract (English)
The annotation of patient organs is a crucial part of various diagnostic and treatment procedures, such as radiotherapy planning. Manual annotation is extremely time-consuming, while its automation using modern image analysis techniques has not yet reached levels sufficient for clinical adoption. This paper investigates the idea of semi-supervised medical image segmentation using human gaze as interactive input for segmentation correction. In particular, we fine-tuned the Segment Anything Model in Medical Images (MedSAM), a public solution that uses various prompt types as additional input for semi-automated segmentation correction. We used human gaze data from reading abdominal images as a prompt for fine-tuning MedSAM. The model was validated on a public WORD database, which consists of 120 CT scans of 16 abdominal organs. The results of the gaze-assisted MedSAM were shown to be superior to the results of the state-of-the-art segmentation models. In particular, the average Dice coefficient for 16 abdominal organs was 85.8%, 86.7%, 81.7%, and 90.5% for nnUNetV2, ResUNet, original MedSAM, and our gaze-assisted MedSAM model, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。